Large language models are already embedded in healthcare in ways most patients never see, from drafting clinical notes during office visits to flagging potential diagnoses that physicians might otherwise miss. The technology is not a distant promise; hospitals and clinics worldwide are deploying these tools right now in documentation, triage, patient communication, and clinical decision support. Yet the picture is far from simple. The same models that outperform doctors on empathy ratings in written exchanges also fabricate medications, recommend contraindicated drugs, and show measurable racial and gender biases. Understanding where medical LLMs genuinely help, where they fail, and where the guardrails remain dangerously thin matters for anyone who interacts with a health system today.
How LLMs Stack Up Against Clinicians in Diagnosis
The most direct question people have about AI in medicine is whether it can actually diagnose disease. The honest answer is that it depends heavily on the task, the model, and the specialty. A systematic review and meta-analysis covering 30 studies found that clinical professionals had higher diagnostic accuracy in about a third of studies, while LLMs (primarily ChatGPT) outperformed clinicians in another third, with the remaining studies showing mixed or comparable results.1PubMed Central. Comparing Diagnostic Accuracy of Clinical Professionals and Large Language Models: Systematic Review and Meta-Analysis That roughly even split tells you something important: these models are not reliably better or worse than doctors. They are better at some things and worse at others.
Rheumatology offers a telling example. When ChatGPT-4 was tested against rheumatologists on real-world case vignettes, it matched their performance on top diagnoses (about 35% versus 39%, a statistically insignificant difference) and actually did slightly better at placing the correct answer somewhere in its top three guesses. Where the model really shone was in detecting inflammatory rheumatic diseases, correctly classifying a higher proportion of those cases than the specialists did. But it paid for that sensitivity with lower specificity, meaning it over-diagnosed inflammatory conditions in patients who did not have them.2PubMed Central. Diagnostic accuracy of a large language model in rheumatology: comparison of physician and ChatGPT-4 That trade-off captures the current state of affairs: LLMs can be useful triage tools that catch things clinicians might miss, but they also generate false positives that need a human to sort out.
The Empathy Surprise
Perhaps the most counterintuitive finding in this field is that AI chatbots consistently produce more empathetic written responses than physicians do. In a study comparing chatbot and physician replies to real patient questions from a public health forum, evaluators rated the chatbot’s responses as higher quality and significantly more empathetic. Nearly 80% of the chatbot’s answers received ratings of “good” or “very good,” compared with about 22% for physician responses. On empathy specifically, physicians’ replies were rated roughly 41% less empathetic, and fewer than 5% of physician answers reached the “empathetic” or “very empathetic” threshold, versus about 45% for the chatbot.3JAMA Internal Medicine. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum
This is not a one-off result. A meta-analysis pooling 13 studies that all used ChatGPT-3.5 or ChatGPT-4 found a large overall effect favoring AI on empathy, roughly equivalent to a two-point increase on a ten-point scale.4PubMed Central. AI chatbots versus human healthcare professionals: a systematic review and meta-analysis of empathy in patient care A separate study focused on cancer-related questions found the same pattern across multiple AI models, with patients rating AI responses as substantially more empathetic than those from oncologists.5npj Digital Medicine. Patient perceptions of empathy in physician and artificial intelligence chatbot responses to patient questions about cancer
Before declaring AI the winner at bedside manner, though, context matters. Physicians answering forum questions are usually doing so quickly, between patients, in short text bursts. The chatbot has no time pressure, no competing demands, and an architecture that tends toward verbose, thorough responses. What the empathy gap really highlights is the crushing time constraints clinicians face rather than some intrinsic emotional intelligence in the software. It also raises a practical opportunity: if AI can handle written patient communication, it could free doctors to spend more of their face-to-face time actually connecting with patients.
Cutting Paperwork With Ambient AI Scribes
If there is one area where LLMs have already crossed from experimental to mainstream in medicine, it is clinical documentation. Ambient AI scribes listen to the doctor-patient conversation and automatically generate draft clinical notes. The appeal is obvious: physicians routinely spend more time on documentation than on direct patient care, and burnout driven by administrative tasks is a leading cause of doctors leaving the profession.
A randomized trial testing two commercial ambient scribes (Nabla and Microsoft’s DAX Copilot) found that both groups of users reported meaningful improvements in well-being scores and reductions in professional burnout. One of the tools also reduced the time clinicians spent writing notes by about 9.5%.6PubMed Central. Ambient AI Scribes in Clinical Practice: A Randomized Trial Another pragmatic trial found that ambient AI reduced time spent on notes by roughly 20 minutes per day, a modest but real recovery of time that would otherwise be spent typing after clinic hours.7PubMed Central. A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being
A separate study found that after 30 days of using an ambient scribe, clinicians reported reduced cognitive load from note-writing, greater ability to focus on patients during visits, and about an hour less documentation done after hours each day.8PubMed Central. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout The technology is not perfect: notes still need physician review, and the time savings are not always as dramatic as vendors claim. But the evidence points clearly in one direction. Ambient scribes reduce the documentation grind, and clinicians who use them feel measurably less burned out.
The Hallucination Problem
The most dangerous failure mode of medical LLMs is hallucination, where the model generates information that is factually wrong, logically inconsistent, or unsupported by clinical evidence. In a healthcare context, a hallucination is not just an embarrassing error. It could mean a fabricated medication, a contraindicated drug recommendation, or a false interpretation of an imaging study.
When researchers assessed LLM-generated summaries of nearly 13,000 sentences from clinical notes, they found hallucinations in about 1.5% of sentences. That sounds low until you consider two things: almost half of those hallucinations were classified as major, meaning they could change a patient’s diagnosis or management if a clinician did not catch the error. The most common type was fabrication, where the model invented clinical details that did not exist in the source material, and these fabrications appeared most frequently in the planning sections of notes, exactly the part that guides what happens to the patient next.9npj Digital Medicine. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation
A broader evaluation tested 11 foundation models across seven medical tasks and found something surprising: general-purpose models like GPT-4 actually produced hallucination-free responses at substantially higher rates than medical-specialized models (a median of about 77% versus 51%). Physician audits of the remaining hallucinations revealed that roughly two-thirds stemmed from failures in causal or temporal reasoning rather than missing knowledge. The model “knew” the relevant facts but assembled them incorrectly. A survey of 70 clinicians across 15 specialties confirmed the real-world stakes: over 90% had personally encountered medical hallucinations from AI, and about 85% believed those hallucinations could cause patient harm.10arXiv. Medical Hallucination in Foundation Models and Their Impact on Healthcare
The implication is sobering. Making a model more “medical” through specialized training does not necessarily make it safer. And the hallucinations that persist are not the kind you can fix by feeding in more medical textbooks. They come from the model’s reasoning architecture itself.
Bias in Diagnosis and Treatment Recommendations
Any tool trained on historical medical data inherits the biases baked into that data, and LLMs are no exception. A systematic review of 24 studies evaluating demographic disparities in medical LLMs found that over 90% identified some form of bias. Gender bias was the most frequently documented, appearing in nearly all studies that looked for it, and racial or ethnic biases showed up in about 91% of the studies that tested for them.11PubMed Central. Evaluating and addressing demographic disparities in medical large language models: a systematic review
The bias is not always subtle. When researchers tested four LLMs on psychiatric case vignettes that included racial characteristics, they found that models frequently proposed different treatment approaches depending on the patient’s race. The effect was strongest in schizophrenia and anxiety cases and was more pronounced when racial information was explicitly stated rather than implied. Treatment recommendations were more biased than diagnostic outputs, meaning the models might arrive at similar diagnoses regardless of race but then recommend different therapies. Among the models tested, some were far worse than others, with one model receiving the maximum bias score for treatment recommendations more often than any competitor.12npj Digital Medicine. Racial bias in AI-mediated psychiatric diagnosis and treatment: a qualitative comparison of four large language models
This is not a theoretical concern for the future. If health systems deploy these tools in clinical decision support without rigorous bias testing, they risk automating and scaling the very disparities that medicine has been trying to address for decades.
Privacy Risks When Patient Data Meets the Cloud
Most commercial LLMs run on cloud infrastructure, which means patient data leaves the hospital’s network. A scoping review of LLM privacy in healthcare found that more than half of the models studied used cloud deployments, and the generative nature of these models creates a specific risk: they can accidentally reproduce private information absorbed during training, including personally identifiable data. The enormous parameter counts and vast training datasets of modern LLMs expand what they memorize and create more potential vectors for data leakage, prompt manipulation, and data-poisoning attacks.13PubMed Central. Considerations for Patient Privacy of Large Language Models in Health Care: Scoping Review
The privacy challenges scale with the deployment environment. A small hospital running a model on a single local server faces different risks than a regional health network sending queries to a third-party cloud service. Both share the core tension: cloud deployment offers cost-effectiveness and scalability that on-premise systems cannot match, but it also introduces risks of data breaches, unauthorized access, and regulatory compliance failures that keep hospital lawyers up at night.14arXiv. SoK: Privacy-aware LLM in Healthcare: Threat Model, Privacy Techniques, Challenges and Recommendations On-premise deployment avoids many of those risks but requires expensive hardware and limits access to the latest model updates. There is no clean answer yet, and the regulatory framework has not caught up with the technology.
Automation Bias and the Oversight Trap
When a doctor receives an AI recommendation, there is a natural tendency to defer to it, even when the doctor’s own assessment was correct. This phenomenon, called automation bias, is one of the least discussed but most consequential risks of integrating LLMs into clinical workflows. In a study of AI-assisted pathology, researchers found that while AI improved overall performance, it also induced a 7% automation bias rate: cases where pathologists had initially made the correct call but then changed their answer to match an incorrect AI suggestion.15Bildverarbeitung für die Medizin. Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology Time pressure made the problem worse, which is unfortunate given that time pressure is the default state in most clinical settings.
The paradox is real: you deploy AI to help busy clinicians make better decisions, but the busier those clinicians are, the more likely they are to uncritically accept whatever the AI says. This dynamic means that the benefits of AI decision support can be partially offset by the very human tendency to trust the machine, and it underscores why “human in the loop” is not an automatic safety guarantee. The human has to actually think independently, which gets harder when you are exhausted and the AI has been right 95% of the time.
Matching Patients to Clinical Trials
One of the more promising and less controversial applications of medical LLMs is matching patients to clinical trials. Today this process is absurdly labor-intensive. A research coordinator might spend 45 minutes or more manually reviewing a single patient’s chart against a trial’s eligibility criteria, and enrollment in cancer trials remains notoriously low despite many patients being eligible.
A scoping review of LLM-based patient-trial matching systems found that most use the model to process eligibility criteria and patient records simultaneously, and several have demonstrated improved accuracy and scalability compared with manual methods.16PubMed Central. Enhancing Patient-Trial Matching With Large Language Models: A Scoping Review of Emerging Applications and Approaches One system, built on open-source LLMs with a retrieval-augmented generation framework, achieved 93% accuracy on a benchmark dataset and 87% in real-world trials. Users reviewing the AI’s eligibility assessments completed their review in under nine minutes per patient on average, representing an 80% time savings over traditional chart review.17PubMed Central. Real-world validation of a multimodal LLM-powered pipeline for high-accuracy clinical trial patient matching Another system, TrialMatchAI, automates the process by ingesting both structured records and unstructured physician notes, working within a lightweight framework suitable for hospital deployment.18Nature Communications. TrialMatchAI: an end-to-end AI-powered clinical trial recommendation system to streamline patient-to-trial matching
The remaining challenges are real but solvable: performance varies across trial types, the systems sometimes struggle when medical records are incomplete, and the matching rationale can be hard for clinicians to interpret. But compared with the diagnostic and treatment-recommendation applications, trial matching carries lower safety risk and higher immediate payoff, making it a natural early adoption target.
LLMs in Low-Resource Health Systems
The potential impact of medical LLMs may be greatest in places with the fewest doctors. A study testing LLMs in a low-resource, multilingual health system found that even when communicating in Kinyarwanda (a language with far less representation in training data than English), the models outperformed local clinicians on clinical reasoning tasks and cost over 500 times less per response.19PubMed Central. Large language models for frontline healthcare support in low-resource settings Separate work has explored fine-tuning open-source LLMs for Vietnamese health communication, demonstrating that even relatively modest adaptation efforts can improve healthcare communication in languages that major commercial models handle poorly.20PubMed. Fine-tuning large language models for improved health communication in low-resource languages
The promise here is significant but comes with caveats. Performance degrades in under-resourced languages, bias testing has been conducted overwhelmingly in English, and the infrastructure requirements for cloud-based models may be hard to meet in remote settings. Still, in a world where billions of people lack access to a single specialist, an imperfect AI that provides reasonable clinical guidance is a meaningful step forward.
Patients Want to Know When AI Is Involved
As these tools proliferate, an important question is whether patients should be told when AI plays a role in their care, and whether that knowledge changes how they feel about the care they receive. A web-based experiment found that when patients learned an AI tool was being used, they perceived information about that tool as more important than, or at least as important as, information about conventional treatment effects. In other words, patients treated the presence of AI as something they deserved to know about, on par with being told about side effects or risks. The study also found that gender, age, and income influenced how important patients considered AI-related disclosures to be, suggesting that consent processes may need to be tailored rather than one-size-fits-all.21PubMed Central. Patient perspectives on informed consent for medical AI: A web-based experiment
Most health systems have not yet formalized consent protocols for AI-assisted care, and the legal landscape is evolving rapidly. Some jurisdictions are beginning to require disclosure, while others treat AI tools the same as any other software in the clinical workflow. The gap between what patients expect and what institutions currently disclose is wide.
Regulation, Liability, and the Accountability Gap
The U.S. Food and Drug Administration has established regulatory guidelines for AI and machine learning medical devices, treating many of them as “Software as a Medical Device” and proposing continuous surveillance requirements to monitor how these tools perform after deployment.22PubMed Central. Advancements in Clinical Evaluation and Regulatory Frameworks for AI-Driven Software as a Medical Device (SaMD) But regulation has not kept pace with the speed at which LLMs are being integrated into clinical workflows. A systematic review of the liability landscape found that there is no single, specific regulation governing who is responsible when an AI-related error harms a patient. The question of whether liability falls on the developer, the hospital, the prescribing physician, or the model itself remains unresolved.23PubMed Central. Defining medical liability when artificial intelligence is applied on diagnostic algorithms: a systematic review
This ambiguity creates practical problems. If an ambient scribe introduces a factual error into a clinical note and a physician signs off on it without catching the mistake, standard malpractice logic suggests the physician bears responsibility. But what if the error was nearly undetectable without re-listening to the entire encounter? What if the model’s training data contained systematic errors that no individual clinician could have anticipated? These questions do not have settled answers, and the legal system is likely to generate them case by case over the coming years rather than through any single legislative act.
Making LLMs Smarter Without Retraining Them
One of the more active areas of development is figuring out how to improve LLM performance in medicine without the expense and rigidity of full model retraining. Two approaches dominate. Fine-tuning adapts the model’s internal parameters to a specific medical domain but is resource-intensive and does not allow for real-time updates as medical knowledge evolves. Retrieval-augmented generation, or RAG, takes a different approach: it leaves the base model untouched but feeds relevant, up-to-date medical information directly into the query. A systematic review found that RAG offers more flexibility and control than fine-tuning, particularly for clinical applications where guidelines change frequently.24PubMed Central. Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines
Ensemble reasoning approaches offer another path forward. Researchers have shown that combining multiple reasoning strategies can improve performance and consistency on medical exam questions, with gains of several percentage points over standard prompting methods on models like GPT-3.5 and Med42.25PubMed Central. Reasoning with large language models for medical question answering These incremental improvements matter because in medicine, a few percentage points of accuracy can translate to thousands of patients correctly diagnosed or correctly routed.
Multimodal Models and the Radiology Frontier
Text-only models are already being surpassed by multimodal LLMs that can process images alongside clinical notes. In radiology, these models combine language understanding with visual analysis, interpreting everything from two-dimensional chest X-rays to three-dimensional CT and MRI scans alongside the patient’s clinical history.26PubMed Central. Multimodal Large Language Models in Medical Imaging: Current State and Future Directions The goal is not to replace radiologists but to integrate visual and textual reasoning in a way that mirrors how a physician actually thinks when reading an image: looking at the scan, reading the clinical question, considering the patient’s history, and synthesizing all of it into a finding.
The technology is still early. Current multimodal medical models lag behind their text-only counterparts in reliability, and the hallucination risks are amplified when the model has to interpret visual data it may not have seen enough of during training. But the trajectory is clear, and radiology, pathology, and dermatology are the specialties most likely to see transformative multimodal AI integration in the near term.
Multi-Agent Systems and What Comes Next
The current generation of medical LLMs operates as single models responding to single prompts. The next generation may look very different. Researchers have begun designing multi-agent AI systems in which several specialized AI agents collaborate on a single clinical problem, each handling a different aspect of care. One proposed architecture for sepsis management envisions seven specialized agents working together: one gathering patient data, another generating a diagnosis, a third recommending treatments, and others managing resources, monitoring the patient, and coordinating communication.27PubMed Central. Multiagent AI Systems in Health Care: Envisioning Next-Generation Intelligence
These systems are largely hypothetical today, but they point toward a future in which AI in healthcare is not a single chatbot answering questions but an integrated network of specialized models, each checking the others’ work. Whether that architecture actually reduces hallucinations and bias or merely distributes them across more components is an open and genuinely difficult question. The engineering is moving faster than the safety science, which is a pattern that should sound familiar to anyone who has watched how healthcare technology tends to roll out.