Artificial intelligence in healthcare faces a web of unresolved problems that span technical, ethical, legal, and economic dimensions. Despite genuine advances in medical imaging, clinical documentation, and drug discovery, most AI tools still struggle to move from controlled research settings into everyday clinical practice. The reasons go well beyond any single flaw: biased training data, opaque decision-making, regulatory gaps, and fragile real-world performance all compound one another. Understanding these hurdles matters not just for technologists but for anyone who will eventually encounter AI-assisted care.
Training Data That Leaves People Out
AI systems learn patterns from the data they are trained on. When that data over-represents certain groups and under-represents others, the resulting tool can perform well for some patients and dangerously poorly for the rest. Several populations have a long history of being absent or misrepresented in the biomedical datasets that feed these algorithms, and the models trained on those datasets tend to reinforce rather than correct those imbalances.1PubMed Central. Addressing bias in big data and AI for health care: A call for open science
Skin-disease detection is a frequently cited example. Diagnostic AI models trained primarily on images of lighter skin tend to underperform on patients with darker skin, simply because the training set did not include enough diversity for the algorithm to learn what pathology looks like across a range of skin tones.2PubMed Central. The Growing Use of Artificial Intelligence in Health Care and Implications for Disparities The same dynamic can play out with age, sex, geography, and socioeconomic status. An algorithm built on data from a handful of well-resourced academic hospitals may not generalize to a rural clinic with a different patient mix, different equipment, or different documentation habits.
The frustrating part is that this problem is well recognized yet difficult to fix at scale. Collecting truly representative datasets requires coordination across institutions, countries, and communities that have historically had good reason to distrust medical research. Retrofitting diversity into datasets after the fact is not straightforward either, because simply rebalancing a dataset can introduce its own distortions.
The Black Box Problem
Many of the most powerful AI models in medicine are deep neural networks, and they share a common trait: no one, including the people who built them, can fully explain how they arrive at a given output. This “black box” characteristic means that when an AI recommends a diagnosis or a treatment plan, patients, physicians, and even the system’s designers cannot trace the reasoning step by step.3Intelligent Medicine. Medical artificial intelligence and the black box problem: a view based on the ethical principle of “do no harm”
This lack of transparency has real consequences for clinical adoption. Physicians are trained to justify their decisions and to be accountable for them. When a black-box algorithm flags a patient as high-risk, the clinician is left in an awkward position: they can see the input data and the output recommendation, but they cannot verify the logic in between. That raises uncomfortable questions about whether a physician can ethically act on a recommendation they cannot scrutinize, and whether they can be held responsible for a diagnosis that came from a system they do not fully understand.4PubMed. Who is afraid of black box algorithms? On the epistemological and ethical basis of trust in medical AI
Researchers have been developing what is often called “explainable AI,” techniques that try to highlight which features or data points drove a particular decision. But progress has been slow, and the gap between deep-learning performance and interpretability remains wide. The insufficient explainability in most existing AI systems is considered one of the main reasons that successful integration into routine clinical practice is still uncommon.5PubMed Central. Unbox the black-box for the medical explainable AI via multi-modal and multi-centre data fusion: A mini-review, two showcases and beyond
Hallucinations and Fabricated Medical Facts
Large language models, the technology behind tools like ChatGPT, are increasingly being explored for clinical documentation, patient communication, and decision support. A central barrier to using them safely is the problem of medical hallucinations: the model generates statements that sound authoritative but are factually wrong or entirely fabricated.6npj Digital Medicine. A multicenter assessment of human oversight of generative AI outputs in simulated clinical decision making
A study evaluating large language models for clinical note generation across nearly 13,000 clinician-reviewed sentences found a hallucination rate of about 1.5%, with roughly 44% of those hallucinations rated as major, meaning they could affect patient diagnosis or management if left uncorrected. The types of errors included fabricated information, negations of what was actually true, and incorrect causal claims. Major hallucinations appeared across all sections of the notes but were most common in treatment plans and clinical assessments.7npj Digital Medicine. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation
A 1.5% error rate might sound small, but in medicine, the denominator matters enormously. Across thousands of clinical notes generated daily in a large hospital system, even a low per-sentence hallucination rate translates into a meaningful number of potentially harmful errors. The only current safeguard is human review, which partially defeats the efficiency gains that motivated using the tool in the first place.
When a Model Leaves the Hospital It Was Trained In
An AI model that performs brilliantly at the institution where it was developed often stumbles when deployed somewhere else. This phenomenon, called domain shift, happens because hospitals differ in their equipment, imaging protocols, patient demographics, coding practices, and even how clinicians write notes. One of the main obstacles to incorporating automated AI decision-making tools in medicine is precisely this failure to generalize across institutions.8PubMed Central. Toward Generalizability in the Deployment of Artificial Intelligence in Radiology: Role of Computation Stress Testing to Overcome Underspecification
Research on chest X-ray classification illustrates how steep the drop can be. In one study, a model trained on data from one population saw accuracy plummet from nearly 80% down to as low as 6.7% when applied to a different population’s images.9Scientific Reports. Addressing cross-population domain shift in chest X-ray classification through supervised adversarial domain adaptation A systematic review of AI tools in radiology found that external validation consistently produced lower performance than internal testing, with specificity often dropping more sharply than sensitivity. In one case, sensitivity barely changed between internal and external validation, but specificity fell from 94% to 70%.10PubMed Central. Assessing the generalizability of artificial intelligence in radiology: a systematic review of performance across different clinical settings
The practical implication is sobering: a tool marketed based on impressive internal results may function much less reliably once it ships to a different hospital across the country, let alone to a clinic in a different part of the world with different equipment. Vendors and health systems are starting to require external validation before deployment, but it is far from standard practice.
Model Drift After Real-World Disruptions
Even a model that performs well at launch can degrade over time as the world changes around it. The COVID-19 pandemic offered a stark demonstration. A population health algorithm used in the Veterans Health Administration to predict readmission and mortality risk saw its performance shift during the pandemic because the patterns in patient data changed dramatically: utilization dropped, new diagnoses surged, and the relationship between lab values and outcomes shifted.11JAMA Health Forum. Understanding Model Drift and Its Impact on Health Care Policy The model had not become buggy; the reality it was built to predict had simply moved out from under it. Any future pandemic, policy change, or major shift in clinical practice could trigger the same kind of drift, and most deployed models lack mechanisms to detect it automatically.
Privacy Risks That Outlive Anonymization
AI in healthcare depends on access to vast amounts of patient data. Sharing that data for model training or research typically involves stripping out names, dates of birth, and other direct identifiers. But de-identification is not as bulletproof as it sounds. Electronic medical records stored and shared across cloud platforms expose patient data to third-party service providers and research collaborators, raising serious concerns about privacy.12Procedia Computer Science. Preserving Privacy of Patients Based on Re-identification Risk
Even after direct identifiers are removed, unique sequences of events in a patient’s record can be enough to re-identify someone. Rare combinations of diagnoses, procedures, and dates can effectively serve as fingerprints, especially when combined with information available from other public sources. These kinds of identifying details are particularly likely to appear in unstructured clinical text, such as physician notes, where a throwaway mention of living arrangements or a specific life event can narrow the pool to a single patient.13PubMed Central. What is the patient re-identification risk from using de-identified clinical free text data for health research? Research has shown that even anonymized text data can be vulnerable to re-identification attacks that effectively reverse the anonymization process.14ACL Anthology. Clinical Text Anonymization, its Influence on Downstream NLP Tasks and the Risk of Re-Identification
Regulation That Cannot Keep Up
Traditional medical device regulation was designed for products that are fixed at the point of approval: a hip implant does not change shape after surgery. AI software, however, can be designed to learn and adapt continuously from new data, which means the product the regulator approved today might behave differently six months from now. Continuously learning algorithms pose significant public health risks if a device can fundamentally alter its own behavior after reaching the market.15American Journal of Law & Medicine. When Medical Devices Have a Mind of Their Own: The Challenges of Regulating Artificial Intelligence
The FDA has acknowledged this gap and begun developing frameworks for regulating adaptive AI. But the challenge goes beyond simply re-testing an updated algorithm on a reference dataset. When humans interact with the tool, any update also changes how clinicians use it, how patients respond, and how workflows accommodate it. A true systems perspective makes regulatory review of updates far more demanding.16npj Digital Medicine. The need for a system view to regulate artificial intelligence/machine learning-based software as medical device Recent analyses have noted that significant weaknesses remain in real-time monitoring, transparency, and bias mitigation even after the FDA’s initial regulatory adjustments.17PubMed Central. The illusion of safety: A report to the FDA on AI healthcare product approvals
Adding to the problem, most AI clinical trials lack the prospective, randomized designs that would generate the strongest evidence for safety and effectiveness. Many registered trials do not report results upon completion, leaving regulators and clinicians with insufficient evidence to assess how well these tools actually perform in practice.18PubMed Central. Characteristics of Artificial Intelligence Clinical Trials in the Field of Healthcare: A Cross-Sectional Study on ClinicalTrials.gov
Who Is Liable When AI Gets It Wrong?
When a physician misdiagnoses a patient, liability frameworks, however imperfect, exist to assign responsibility. When an AI system contributes to a wrong diagnosis, the picture fragments. Was the error caused by a flaw in the developer’s programming? A defect introduced by the manufacturer? A self-learning algorithm that drifted into an erroneous pattern on its own? Or did the patient provide misleading input? Any of these, or a combination, could be the cause.19Humanities and Social Sciences Communications. Civil liability for the actions of autonomous AI in healthcare: an invitation to further contemplation
The current regulatory framework for medical liability when AI is involved has been described as inadequate and requiring urgent intervention. No single regulation governs the liability of the various parties in the AI supply chain, from data providers and algorithm developers to hospitals and clinicians who use the tool at the bedside.20PubMed Central. Defining medical liability when artificial intelligence is applied on diagnostic algorithms: a systematic review Until clearer rules emerge, hospitals deploying AI tools and physicians relying on them operate in a legal gray zone that creates uncertainty for everyone, including patients seeking recourse after a bad outcome.
Alert Fatigue and Workflow Friction
Even when an AI tool works as intended, its clinical value depends on how well it fits into the messy reality of a hospital workflow. Clinical decision support alerts embedded in electronic health records are a telling case study. An analysis of physician discourse found that the most frequently discussed problems with these alerts were burnout and fatigue, reported in over 40% of the discussions, followed by alert overwhelm at nearly 39%.21Applied Clinical Informatics. Systemic Challenges in EHR Alert Design: A Thematic Analysis of Physician Discourse on X (Formerly Twitter) Clinicians swamped with irrelevant or low-value notifications learn to click past them reflexively, which means even genuinely important AI-generated alerts can go unheeded.
Research on AI-enhanced health record systems has found that while automated documentation and decision support can reduce administrative workload by a meaningful amount, these gains come with trade-offs including communication fragmentation between nurses and physicians, cognitive overload, and excessive reliance on electronic intermediaries rather than direct conversation.22The Review of Diabetic Studies. Evaluating How Artificial Intelligence And Electronic Health Record Systems Influence Physician–Nurse Communication, Workflow Efficiency, And Clinical Decision-Making The technology can save time with one hand and steal it with the other if the design does not account for how clinical teams actually communicate.
Automation Bias and Over-Reliance
There is a well-documented human tendency to defer to automated systems, especially under time pressure, and clinical settings are full of time pressure. This automation bias means clinicians may accept an AI recommendation without exercising the independent judgment the tool was designed to augment rather than replace.23Journal of Safety Science and Resilience. Exploring the risks of automation bias in healthcare artificial intelligence applications: A Bowtie analysis The risk is circular: the more accurate a tool is most of the time, the more clinicians trust it, and the less likely they are to catch the cases where it fails. Training programs and interface designs that encourage critical evaluation rather than passive acceptance remain underdeveloped.
Security Vulnerabilities
AI systems in medicine face a threat that traditional medical devices do not: adversarial attacks. Researchers have demonstrated that small, carefully crafted changes to medical images, often invisible to the human eye, can cause AI systems to misclassify findings in ophthalmology, radiology, and pathology.24PubMed. Adversarial attack vulnerability of medical image analysis systems: Unexplored factors The financial incentives in healthcare, from insurance fraud to manipulating diagnostic outcomes, make this more than a theoretical concern. As AI tools become more embedded in clinical workflows, the attack surface grows.
The Cost Problem Nobody Talks About
Deploying healthcare AI is expensive, and the true costs are poorly understood. Most economic evaluations of AI in health focus on the price of the software or its effect on downstream spending but ignore lifecycle costs like development, validation, integration into existing systems, ongoing maintenance, retraining as data shifts, and eventual decommissioning. Energy consumption, data hosting, and cloud infrastructure costs are almost never factored in.25medRxiv. Costing Methods for Artificial Intelligence: Systematic Review and Recommended Cost Inventory for in Health Technology Assessment Large healthcare systems exploring generative AI for operational tasks have already encountered significant challenges with computational costs and model reliability.26PubMed Central. Generative AI costs in large healthcare systems, an example in revenue cycle
This has implications for equity. While AI adoption accelerates in wealthy health systems, its potential to improve care in low-resource settings remains under-realized. Shortages of health workers and underdeveloped infrastructure in low- and middle-income countries already hinder the delivery of equitable care, and the high costs and technical demands of AI tools risk widening the gap rather than narrowing it.27PubMed Central. Designing AI tools to advance health equity in resource-constrained low- and middle-income countries
Data Silos and Interoperability
Healthcare data is notoriously fragmented. Different hospitals use different electronic health record systems, different coding standards, and different data formats. Efforts to standardize data exchange through frameworks like FHIR (Fast Healthcare Interoperability Resources) have made progress, but the heterogeneous structures and formats of health data remain a significant barrier, particularly when critical information is buried in unstructured text rather than neatly coded fields.28PubMed Central. FHIR-GPT Enhances Health Interoperability with Large Language Models An AI model that needs comprehensive patient data to make good predictions can only be as good as the data pipeline feeding it, and those pipelines are often leaky and inconsistent.
Who Owns the Data?
As AI systems require more patient data to train and improve, a thorny question has grown louder: who actually owns medical data? Current laws in both the United States and the European Union provide privacy protections but do not explicitly establish ownership rights. Patients, clinicians, researchers, device manufacturers, and institutions can all stake plausible claims, and no specific legislation clearly arbitrates among them.29PubMed Central. The need for clear medical data ownership laws Without resolved ownership, basic questions about consent, compensation, and control over how data is used for commercial AI products remain unanswered. This ambiguity slows data sharing for legitimate research and creates openings for exploitation.
The Patient-Physician Relationship
Beyond the technical obstacles, there is a distinctly human concern. Qualitative research involving multiple stakeholders has found that AI could alter fundamental aspects of the patient-physician relationship. On one hand, AI can free up clinician time by handling administrative burdens, potentially putting the patient back at the center of the caring process. On the other, there is a real risk that leaning on algorithmic outputs erodes the holistic approach to care, the part where a physician reads a patient’s anxiety, picks up on an unspoken concern, or factors in a life circumstance that no dataset captures.30SAGE Journals / Digital Health. Critical analysis of the AI impact on the patient-physician relationship: A multi-stakeholder qualitative study The worry is not that AI will replace physicians outright, but that it will subtly reshape clinical encounters in ways that make care feel more transactional and less personal.
Environmental Costs of Healthcare AI
There is an irony in developing AI tools that aim to improve health while contributing to environmental harms that undermine it. The energy demands of training and running large AI models are substantial, and healthcare AI is increasingly reliant on technologies like triage algorithms, electronic patient records, and surgical robotics that require continuous computation. These systems contribute to carbon emissions through the electricity needed for data centers, the manufacturing of specialized hardware, and the cooling infrastructure to keep it running.31Health and Technology. Environmental impacts of artificial intelligence in health care: considerations and recommendations The environmental costs span the entire AI lifecycle, from data collection and model training to deployment and maintenance, and they are rarely accounted for in discussions of AI’s promise for medicine.32PubMed Central. The Environmental Costs of Artificial Intelligence for Healthcare
Rare Diseases and Pediatric Gaps
AI tends to perform best where data is abundant, which means it tends to perform worst precisely where it could be most needed. Children with rare genetic diseases, for instance, sometimes present with distinctive facial features that a well-trained model could theoretically flag for early diagnosis. But building such models is exceptionally difficult because of extreme data scarcity, strict privacy constraints in pediatric settings, and limited data sharing among the small number of institutions that treat these conditions.33arXiv. Synthetic Data Alone is Enough? Rethinking Data Scarcity in Pediatric Rare Disease Recognition Researchers have explored generating synthetic training images to fill the gap, but validating those approaches is itself constrained by the same scarcity problem. It is a catch-22 that illustrates a broader truth about healthcare AI: the populations with the fewest existing resources often have the least to gain from tools built on majority data.