Natural language processing applied to clinical notes is rapidly reshaping how health systems extract meaning from the enormous volume of free text buried in electronic health records. Most of the clinically rich detail about a patient’s condition, reasoning, and social context lives not in structured drop-down fields or billing codes, but in the narrative text that clinicians type or dictate during encounters. A feasibility study across millions of patient records found that only about 13% of the concepts extracted from free text overlapped with the structured data in the same records, meaning the vast majority of what clinicians write down never makes it into the searchable, computable parts of the chart.1PubMed Central. Using Structured Codes and Free-Text Notes to Measure Information Complementarity in Electronic Health Records: Feasibility and Validation Study Unlocking that text with NLP promises everything from faster drug-safety surveillance to more equitable care, but the technology comes with real limitations that clinicians and patients should understand.
Why So Much Valuable Information Lives in Free Text
When a physician sees a patient, the resulting note typically includes the reason for the visit, a narrative history, physical exam findings, assessment, and a plan. Structured fields capture diagnosis codes, lab values, and medication orders, but the reasoning behind those decisions, the nuances of symptom descriptions, and references to a patient’s living situation or substance use tend to appear only in prose. That same feasibility study showed that structured data is more likely to be echoed somewhere in the unstructured notes than the reverse: about 42% of structured concepts had a matching concept in the text, while only 13% of text-derived concepts appeared in structured form.1PubMed Central. Using Structured Codes and Free-Text Notes to Measure Information Complementarity in Electronic Health Records: Feasibility and Validation Study In practical terms, if you only look at the coded data, you are missing the majority of the clinical picture.
From Dictionaries to Deep Learning
Early clinical NLP systems relied on hand-built dictionaries and rules. A developer would manually encode patterns like “no evidence of” or “denies” to flag negated findings, and maintain concept lists linking shorthand terms to standardized medical vocabularies. These rule-based pipelines could handle rare concepts and allowed case-by-case error correction, but they were brittle and expensive to maintain.2PubMed Central. Ensembles of Natural Language Processing Systems for Portable Phenotyping Solutions As de-identified clinical text became more widely available to the research community, machine learning approaches, and later deep learning, overtook these older systems in accuracy for tasks like extracting medical concepts from notes.2PubMed Central. Ensembles of Natural Language Processing Systems for Portable Phenotyping Solutions
The real leap came with transformer-based models. General-purpose models like BERT were adapted for biomedicine (BioBERT) and for clinical text (ClinicalBERT), but these were relatively small. The GatorTron project pushed clinical language models much further, training from scratch on over 90 billion words of text, including more than 82 billion words of de-identified clinical notes.3arXiv. GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records Scaling the model from 110 million parameters to 8.9 billion improved performance across five clinical NLP tasks, including nearly 10% gains in accuracy for natural language inference and medical question answering.3arXiv. GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records A later generative version, GatorTronGPT, outperformed existing biomedical models on relation extraction tasks by 3 to 10% compared with the next-best model, and showed consistent improvement as model size increased.4PubMed Central. A study of generative large language model for medical research and healthcare
Why Local Fine-Tuning Matters
A model trained on clinical notes from one institution or country does not automatically perform well on notes from another. Vocabulary, abbreviation habits, documentation templates, and patient demographics all vary. A study evaluating clinical large language models on electronic health records from South Asia found that without local fine-tuning, performance dropped by at least 15% on concept extraction and 35% on question-answering tasks when tested on notes from the local population.5PubMed Central. Evaluating large language models for clinical note processing: local fine-tuning and internal-external validation using electronic health records from South Asia Fine-tuning with local data recovered much of that loss, yielding 7.5% to 15% improvements on concept extraction and 27% to 53% on question answering.5PubMed Central. Evaluating large language models for clinical note processing: local fine-tuning and internal-external validation using electronic health records from South Asia One exception was ChatGPT, which performed better on the local dataset than on the public one even without fine-tuning, suggesting that very large general-purpose models may generalize better out of the box.5PubMed Central. Evaluating large language models for clinical note processing: local fine-tuning and internal-external validation using electronic health records from South Asia The practical takeaway is that deploying clinical NLP at a new site without adaptation is risky.
Decoding What Clinicians Actually Mean
Clinical language is full of traps for automated systems. A note that says “no signs of pneumonia” contains the word pneumonia, but the meaning is the opposite of what a naive keyword search would flag. Negation detection has been a core challenge in clinical NLP for decades. A combined pipeline using the NegEx algorithm and a convolutional neural network achieved an F1 score of 0.94 for affirmative statements and 0.85 for negated ones, and held up when tested on notes from a completely different hospital system.6PubMed Central. Negation recognition in clinical natural language processing using a combination of the NegEx algorithm and a convolutional neural network That same pipeline also achieved near-perfect scores on Spanish negation benchmarks, showing it can transfer across languages with appropriate training data.6PubMed Central. Negation recognition in clinical natural language processing using a combination of the NegEx algorithm and a convolutional neural network
Beyond negation, clinicians use speculation (“possible cellulitis”), temporal references (“history of stroke”), abbreviations that differ across specialties, and hedging language that all need to be interpreted correctly before a system can reliably determine what is actually happening with a patient.
Identifying Patient Cohorts from Narratives
One of the most established uses of clinical NLP is computational phenotyping, or figuring out which patients in a health system have a given condition based on their records. This matters for clinical trials, quality measurement, pharmacogenomics, and large-scale genetic association studies.7PubMed Central. Natural Language Processing for EHR-Based Computational Phenotyping The problem is that diagnosis codes alone are unreliable. A billing code for diabetes might be applied for a screening visit where diabetes was ruled out. The most relevant evidence for whether a patient truly has a condition often exists only in narrative text.8PubMed Central. Comparing deep learning and concept extraction based methods for patient phenotyping from clinical narratives Deep learning approaches like convolutional neural networks have shown they can match or exceed traditional concept-extraction pipelines for this kind of patient classification, and they require less manual feature engineering.8PubMed Central. Comparing deep learning and concept extraction based methods for patient phenotyping from clinical narratives
Automating Medical Coding
After every hospital discharge, coders read the clinical notes and assign ICD codes that determine how the hospital gets paid, how diseases are tracked, and how insurance claims are processed. This work is slow, expensive, and error-prone. NLP systems trained on discharge summaries can assist by suggesting codes automatically.9arXiv. A Systematic Literature Review of Automated ICD Coding and Classification Systems using Discharge Summaries Recent work using fine-tuned large language models has outperformed standard approaches on benchmark datasets for ICD coding, with particularly strong improvements from encoder-decoder model architectures.10PubMed. How to leverage large language models for automatic ICD coding Full automation is still out of reach because of the sheer number of possible codes and the consequences of getting them wrong, but even partial automation can reduce turnaround time and catch codes that human coders miss.
Spotting Adverse Drug Events
Drug side effects are heavily underreported in structured fields. A clinician might note in free text that a patient developed a rash after starting a new antibiotic, but that observation may never be recorded as a formal adverse event. NLP enables post-marketing drug safety surveillance by mining clinical notes for mentions of drugs, symptoms, and the relationships between them.11Journal of Biomedical Informatics. Identifying adverse drug event information in clinical notes with distributional semantic representations of context A scoping review of supervised learning methods for adverse drug event detection found that named entity recognition and relation extraction were the most common tasks, with LSTM networks and conditional random fields as the most frequently used methods.12PLoS ONE. Adverse drug event detection using natural language processing: A scoping review of supervised learning methods More recent work using transformer-based models has shown strong performance, with architectures like Clinical-Longformer proving practical for processing the long documents that real clinical notes tend to be.13PLoS One. Developing a natural language processing system using transformer-based models for adverse drug event detection in electronic health records
Surfacing Social Determinants of Health
A patient’s housing situation, employment status, food security, and substance use can influence health outcomes as much as any medication. These details are often mentioned in clinical notes but almost never captured in structured fields. NLP provides a way to systematically extract social determinants of health from narrative text, though most prior efforts have been limited to single institutions or narrow patient populations.14npj Digital Medicine. Social determinants of health extraction from clinical notes across institutions using large language models Cross-institutional work using large language models is beginning to address generalizability, building corpora from multiple sites and testing whether models trained at one hospital can reliably detect social factors at another.14npj Digital Medicine. Social determinants of health extraction from clinical notes across institutions using large language models This kind of data, once extracted at scale, could feed into population health programs that target upstream drivers of poor outcomes rather than just treating the downstream medical consequences.
Ambient AI Scribes and the Documentation Burden
Clinician burnout is a well-documented crisis, and documentation is one of its biggest drivers. Ambient AI scribes, which use speech recognition and NLP to listen to patient encounters and automatically draft clinical notes, represent one of the most visible near-term applications of this technology. A study in JAMA Network Open found that after 30 days using an ambient AI scribe, the share of clinicians experiencing burnout dropped from about 52% to 39%.15PubMed Central. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout Clinicians also reported spending nearly an hour less per day on after-hours documentation and feeling better able to give patients their undivided attention during visits.15PubMed Central. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout These tools are already deployed at major health systems, making them the NLP application that most patients are likely to encounter in the near future.
Protecting Patient Privacy at Scale
Sharing clinical notes for research requires stripping out protected health information: names, dates, addresses, medical record numbers, and dozens of other identifiers. Doing this manually is painfully slow. Automated NLP-based de-identification pipelines can process notes at enterprise scale, removing or replacing identifiers before the text reaches researchers.16PubMed Central. Natural Language Processing for Enterprise-scale De-identification of Protected Health Information in Clinical Notes A systematic review of de-identification approaches for clinical text confirmed that NLP is now the standard method for automated removal of sensitive information from medical datasets.17PubMed. De-identification of clinical data: A systematic review of free text, image and tabular data approaches Still, the downstream effects of de-identification on clinical utility need attention. A study within the Veterans Health Administration found that automated de-identification can sometimes remove or distort clinically meaningful content, particularly when an identifier (like a facility name) also carries diagnostic information.18PubMed. Text de-identification for privacy protection: a study of its impact on clinical text information content
When Models Hallucinate
The most serious safety concern with large language models in clinical settings is hallucination, where the model produces text that sounds plausible but is factually wrong. A framework for assessing hallucination rates in clinical note summarization examined nearly 13,000 sentences across 450 notes and identified 191 hallucinated sentences, a rate of about 1.5%.19PubMed Central. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation Of those, 44% were classified as major hallucinations that could affect diagnosis or management if left uncorrected. The most common type was fabrication, which appeared most often in the planning section of notes.19PubMed Central. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation A scoping review found that hallucination rates vary wildly depending on the task, spanning from under 1% in closely monitored deployments to over 30% in ambient transcription and zero-shot extraction settings.20SAIMSARA Journal. LLM Hallucinations in Clinical Documentation: Scoping Review with ☸️SAIMSARA The consensus from the field is that these tools should function as assistive systems within human-in-the-loop workflows rather than autonomous documentation agents.20SAIMSARA Journal. LLM Hallucinations in Clinical Documentation: Scoping Review with ☸️SAIMSARA
Bias Embedded in Clinical Language
NLP does not just passively read clinical notes. It also inherits and can amplify the biases contained in them. A study of emergency psychiatry notes developed a “negative sentiment ratio” measuring the density of negative language used to describe patients. That ratio turned out to be the strongest predictor of a schizophrenia diagnosis in the model, and it interacted significantly with patient race: a high negative sentiment ratio increased the odds of a schizophrenia diagnosis more for patients identified as Black, Hispanic or Latino, or “Some Other Race” than for white patients.21arXiv. Bias Detection in Emergency Psychiatry: Linking Negative Language to Diagnostic Disparities The implication is sobering. If an NLP system is trained on notes where certain patients are described in more stigmatizing language, it will learn to associate that language pattern with particular diagnoses, potentially reinforcing diagnostic disparities at scale. Any clinical NLP deployment needs active monitoring for these patterns, not just accuracy metrics.
Plugging NLP Into Health IT Standards
Extracting insights from notes is only useful if those insights can flow back into the systems clinicians actually use. The dominant standard for health data exchange is HL7 FHIR, and work is underway to bridge NLP outputs with FHIR resources. Researchers at Mayo Clinic implemented a clinical NLP pipeline enhanced with an FHIR-based type system that could transform medication mentions extracted from free text into structured FHIR resources, making them available for clinical decision support and research queries alongside data from structured fields.22PubMed Central. Integrating Structured and Unstructured EHR Data Using an FHIR-based Type System: A Case Study with Medication Data A related effort developed NLP extensions for Clinical Quality Language, the standard used to define quality measures and phenotyping algorithms, so that NLP-derived data could be incorporated directly into computable clinical quality measures.23PubMed Central. CQL4NLP: Development and Integration of FHIR NLP Extensions in Clinical Quality Language for EHR-driven Phenotyping Without this kind of standards integration, NLP remains a research tool sitting outside the clinical workflow.
Combining Notes with Images and Time-Series Data
Text from clinical notes is just one modality. Patients also generate lab values over time, imaging studies, vital sign trends, and structured demographic data. Multimodal AI models that combine all of these are beginning to outperform single-source approaches. A framework tested on over 34,000 samples from the MIMIC database trained more than 14,000 models using every possible combination of four data types: tabular, time-series, text, and images. The multimodal models consistently outperformed single-source models by 6 to 33% across tasks including chest pathology diagnosis, length-of-stay prediction, and 48-hour mortality prediction.24npj Digital Medicine. Integrated multimodal artificial intelligence framework for healthcare applications Clinical notes contributed meaningfully to these gains because they captured context and clinical reasoning that no amount of structured data alone could replicate.
Synthetic Clinical Notes as a Privacy Workaround
Access to real clinical notes is tightly restricted for good reason, but this limits NLP research. Synthetic clinical text, generated by language models trained on real de-identified records, offers a workaround. A systematic review found that synthetic note generation has grown rapidly since 2018, with applications spanning data augmentation, corpus building, privacy preservation, and annotation support.25arXiv. Generation of Synthetic Clinical Text: A Systematic Review Early experiments showed that NLP models trained on synthetic notes approached the performance of models trained on real ones for some tasks, though a gap remained.26ACL Anthology. Towards Automatic Generation of Shareable Synthetic Clinical Notes Using Neural Language Models A more recent hybrid approach combines de-identification with LLM-generated text, retaining 36 to 61% of the original note’s content and filling in the rest with synthetic language to maintain information coverage while reducing re-identification risk.27PubMed Central. Not Fully Synthetic: LLM-based Hybrid Approaches Towards Privacy-Preserving Clinical Note Sharing
The Multilingual Gap
Most clinical NLP development has happened in English, which means hospitals and health systems operating in other languages are largely left behind. A cross-lingual analysis noted that available resources are heavily biased toward English, restricting development of medical NLP models in other languages and limiting their adoption in healthcare AI globally.28PubMed. Domain-Specific Multilingual Strategies for Medical NLP: A Cross-Lingual Analysis of Orthographic and Phonemic Representations A survey of medical NLP across multiple languages confirmed that annotated datasets and language-specific models remain scarce, particularly for low-resource languages.29PubMed Central. Exploring the Latest Highlights in Medical Natural Language Processing across Multiple Languages: A Survey This is not just an academic concern. Billions of people receive care documented in languages for which no reliable clinical NLP tools exist, meaning the potential benefits of this technology are concentrated in English-speaking health systems for now.
Making Notes Understandable to Patients
Since 2021, patients in the United States have had the legal right to access their clinical notes through patient portals. The problem is that these notes were written for other clinicians, not for patients. They are dense with abbreviations, jargon, and implicit reasoning that most people cannot parse. NLP-based text simplification offers a path toward making these notes genuinely useful to the people they describe, automatically translating medical language into plain terms while preserving clinical accuracy.30PubMed Central. Towards more patient friendly clinical notes through language models and ontologies Even in ophthalmology, where terminology is unusually specialized, modern LLMs have outperformed older BERT-based models at identifying medication information from progress notes, suggesting the tools are becoming flexible enough to handle the most jargon-heavy specialties.31AMIA Annual Symposium Proceedings. Evaluating the Performance of Large Language Models for Named Entity Recognition in Ophthalmology Clinical Free-Text Notes If done well, this could turn open notes from a source of patient anxiety into a genuine tool for shared decision-making.