Faculty, Staff and Student Publications

Language

English

Publication Date

11-1-2025

Journal

Journal of Biomedical Informatics

DOI

10.1016/j.jbi.2025.104930

PMID

41138954

PMCID

PMC13036512

PubMedCentral® Posted Date

4-1-2026

PubMedCentral® Full Text Version

Author MSS

Abstract

Objective: Multimodal large language models (LLMs) offer new potential for enhancing cardiovascular decision support, particularly in interpreting echocardiographic data. This study systematically evaluates and benchmarks foundation models from diverse domains on echocardiogram-based tasks to assess their effectiveness, limitations and potential in clinical cardiovascular applications.

Methods: We curated three cardiovascular imaging datasets-EchoNet-Dynamic, TMED2, and an expert-annotated echocardiogram (TTE) dataset-to evaluate performance on four critical tasks: (1) cardiac function evaluation through ejection fraction (EF) prediction, (2) cardiac view classification, (3) aortic stenosis (AS) severity assessment, and (4) cardiovascular disease classification. We evaluated six multimodal LLMs: EchoClip (cardiovascular-specific), BiomedGPT and LLaVA-Med (medical-domain), and MiniCPM-V 2.6, LLaMA-3-Vision-Alpha, and Gemini-1.5 (general-domain). Models were assessed using zero-shot, few-shot, and fine-tuning strategies, where applicable. Performance was measured using mean absolute error (MAE) and root mean squared error (RMSE) for EF prediction, and accuracy, precision, recall, and F1 score for classification tasks.

Results: Domain-specific models such as EchoClip demonstrated the strongest zero-shot performance in EF prediction, achieving an MAE of 10.34. General-domain models showed limited effectiveness without adaptation, with MiniCPM-V 2.6 reporting an MAE of 251.92. Fine-tuning significantly improved outcomes; for example, MiniCPM-V 2.6's MAE decreased to 31.93, and view classification accuracy increased from 20 % to 63.05 %. In classification tasks, EchoClip achieved F1 scores of 0.2716 for AS severity and 0.4919 for disease classification but exhibited limited performance in view classification (F1 = 0.1457). Few-shot learning yielded modest gains but was generally less effective than fine-tuning.

Conclusions: This evaluation and benchmarking study demonstrated the importance of domain-specific pretraining and model adaptation in cardiovascular decision support tasks. Cardiovascular-focused models and fine-tuned general-domain models achieved superior performance, especially for complex assessments such as EF estimation. These findings offer critical insights into the current capabilities and future directions for clinically meaningful AI integration in cardiovascular medicine.

Keywords

Humans, Echocardiography, Cardiovascular Diseases, Decision Support Systems, Clinical, Large Language Models, Cardiovascular Disease, Multimodal Large Language Models, Transthoracic Echocardiogram (TTE), Fine-tuning, AI

Published Open-Access

yes

Share

COinS
 
 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.