Faculty, Staff and Student Publications
Language
English
Publication Date
5-5-2026
Journal
Communications Medicine
DOI
10.1038/s43856-026-01621-7
PMID
42086912
PMCID
PMC13350886
PubMedCentral® Posted Date
5-5-2026
PubMedCentral® Full Text Version
Post-print
Abstract
Background: Long COVID affects a substantial proportion of the over 778 million individuals infected with SARS-CoV-2, yet predictive models remain limited in scope. While existing efforts, such as the National COVID Cohort Collaborative (N3C), have leveraged electronic health record (EHR) data for risk prediction and identification, accumulating evidence points to additional contributions from social, behavioral, and genetic factors.
Methods: Using a diverse cohort of SARS-CoV-2-infected individuals (n > 17,200) from the NIH All of Us Research Program, we investigated whether integrating EHR data with survey-based and genomic information improves model performance.
Results: Our multi-scale approach outperforms EHR-only model's area under the receiver operating curve 0.736 (95% CI: 0.730, 0.741), achieving an area of 0.748 (0.741,0.755). Among the top predictors, active-duty service status, and self-reported fatigue are the most informative survey features.
Conclusions: These findings highlight the importance of incorporating multi-scale data to improve risk stratification and inform personalized interventions for long COVID. However the relative increase in accuracy is modest, and the cost of collecting genetic and survey data should be considered before implementation.
Keywords
Viral infection, Predictive markers
Published Open-Access
yes
Recommended Citation
Guardo, Christopher; Xinmeng, Zhang; Gangireddy, Srushti; et al., "Multi-scale Data Improves Performance of Machine Learning Model for Long COVID Identification" (2026). Faculty, Staff and Student Publications. 1000.
https://digitalcommons.library.tmc.edu/uthshis_docs/1000