Faculty, Staff and Student Publications
Language
English
Publication Date
1-1-2025
Journal
AMIA Annual Symposium Proceedings
PMID
40502248
PMCID
PMC12150740
PubMedCentral® Posted Date
6-10-2025
PubMedCentral® Full Text Version
Post-print
Abstract
The progress in natural language processing (NLP) using large language models (LLMs) has greatly improved patient information extraction from clinical narratives. However, most methods based on the fine-tuning strategy have limited transfer learning ability for cross-domain applications. This study proposed a novel approach that employs a soft prompt-based learning architecture, which introduces trainable prompts to guide LLMs toward desired outputs. We examined two types of LLM architectures, including encoder-only GatorTron and decoder-only GatorTronGPT, and evaluated their performance for the extraction of social determinants of health (SDoH) using a cross-institution dataset from the 2022 n2c2 challenge and a cross-disease dataset from the University of Florida (UF) Health. The results show that decoder-only LLMs with prompt tuning achieved better performance in cross-domain applications. GatorTronGPT achieved the best F1 scores for both datasets, outperforming traditional fine-tuned GatorTron by 8.9% and 21.8% in a cross-institution setting, and 5.5% and 14.5% in a cross-disease setting.
Published Open-Access
yes
Recommended Citation
Peng, Cheng; Yu, Zehao; Smith, Kaleb E; et al., "Enhancing Cross-Domain Generalizability in Social Determinants of Health Extraction with Prompt-Tuning Large Language Models" (2025). Faculty, Staff and Student Publications. 944.
https://digitalcommons.library.tmc.edu/uthshis_docs/944