Faculty, Staff and Student Publications
Language
English
Publication Date
7-3-2026
Journal
Briefings in Bioinformatics
DOI
10.1093/bib/bbag367
PMID
42430786
PMCID
PMC13354062
PubMedCentral® Posted Date
7-10-2026
PubMedCentral® Full Text Version
Post-print
Abstract
Large language models (LLMs) are deep learning-based artificial intelligence models that have achieved remarkable success in natural language processing. Typically composed of neural networks with billions of parameters, they are trained on massive unlabeled datasets using self-supervised or semi-supervised learning. Beyond language, LLMs hold immense potential for addressing complex bioinformatics challenges. This review provides a comprehensive overview of transformer-based model applications in genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. We discuss critical components, including tokenization strategies for diverse biological data, transformer architectures, attention mechanisms, and pretraining approaches. We also survey currently available foundation models and their downstream applications across bioinformatics domains. Finally, we highlight major challenges that remain insufficiently addressed in prior reviews and outline future perspectives and design principles for next-generation biological language models, offering practical guidance for both users and developers.
Keywords
Large Language Models, Computational Biology, Humans, Deep Learning, Genomics, Natural Language Processing, Proteomics, Drug Discovery, Neural Networks, Computer, Single-Cell Analysis, language model, foundation model, transformer architecture, multi-omics application, drug discovery, single-cell analysis
Published Open-Access
yes
Recommended Citation
Liu, Jiajia; Yang, Mengyuan; Yu, Yankai; et al., "Advancing Bioinformatics With Language Models: Components, Applications, and Perspectives" (2026). Faculty, Staff and Student Publications. 978.
https://digitalcommons.library.tmc.edu/uthshis_docs/978