Skip to main navigation Skip to search Skip to main content

Phenotype Discovery and Geographic Disparities of Late-Stage Breast Cancer Diagnosis across U.S. Counties: A Machine Learning Approach

Weichuan Dong, Wyatt P. Bensken, Uriel Kim, Johnie Rose, Nathan A. Berger, Siran M. Koroukian

Research output: Contribution to journalArticlepeer-review

Abstract

Background: Disparities in the stage at diagnosis for breast cancer have been independently associated with various contextual characteristics. Understanding which combinations of these characteristics indicate highest risk, and where they are located, is critical to targeting interventions and improving outcomes for patients with breast cancer. Methods: The study included women diagnosed with invasive breast cancer between 2009 and 2018 from 680 U.S. counties participating in the Surveillance, Epidemiology, and End Results program. We used a machine learning approach called Classification and Regression Tree (CART) to identify county "phenotypes," combinations of characteristics that predict the percentage of patients with breast cancer presenting with late-stage disease. We then mapped the phenotypes and compared their geographic distributions. These findings were further validated using an alternate machine learning approach called random forest. Results: We discovered seven phenotypes of late-stage breast cancer. Common to most phenotypes associated with high risk of late-stage diagnosis were high uninsured rate, low mammography use, high area deprivation, rurality, and high poverty. Geographically, these phenotypes were most prevalent in southern and western states, while phenotypes associated with lower percentages of late-stage diagnosis were most prevalent in the northeastern states and select metropolitan areas. Conclusions: The use of machine learning methods of CART and random forest together with geographic methods offers a promising avenue for future disparities research.

Original languageEnglish (US)
Pages (from-to)66-76
Number of pages11
JournalCancer Epidemiology Biomarkers and Prevention
Volume31
Issue number1
DOIs
StatePublished - Jan 2022

ASJC Scopus subject areas

  • General Medicine

Fingerprint

Dive into the research topics of 'Phenotype Discovery and Geographic Disparities of Late-Stage Breast Cancer Diagnosis across U.S. Counties: A Machine Learning Approach'. Together they form a unique fingerprint.

Cite this