AI Bias in Healthcare Examples: Impacts and Ethical Issues

AI systems used in healthcare have already produced documented, measurable harm along lines of race, gender, age, and geography. The most striking cases are not hypothetical: a widely used algorithm steered Black patients away from extra medical care they needed, skin-cancer classifiers performed far worse on darker skin, and language models absorbed racial stereotypes straight from clinical notes. These are not isolated glitches but patterns rooted in how the data was collected, who it represents, and what the algorithms were built to optimize.

The Algorithm That Mistook Cost for Illness

One of the most cited examples of AI bias in healthcare involves a commercial risk-prediction algorithm used to manage the care of roughly 200 million patients in the United States each year. A 2019 analysis published in Science found that the algorithm exhibited significant racial bias: at any given risk score, Black patients were considerably sicker than White patients with the same score. The gap was not subtle. Black patients flagged at the same risk level showed more signs of uncontrolled chronic conditions like diabetes and hypertension than their White counterparts.1PubMed. Dissecting racial bias in an algorithm used to manage the health of populations

The root cause turned out to be the choice of what the algorithm predicted. Instead of predicting how sick someone was, it predicted how much money would be spent on their care in the coming year. Because of longstanding inequities in access to healthcare, less money is typically spent on Black patients than on White patients with the same conditions. The algorithm learned that pattern and treated it as a signal of lower need. The result was that Black patients had to be significantly sicker before the system flagged them for extra attention. This case became a touchstone in the field because the bias was not introduced by a programmer with prejudiced intentions. It emerged from a seemingly neutral design choice interacting with a deeply unequal system.

Skin Cancer Classifiers and the Color of the Data

Dermatology has become one of the most active testing grounds for AI diagnostics, and one of the starkest illustrations of what happens when training data skews toward one population. AI models for detecting skin cancer and other conditions consistently perform worse on darker skin tones. A narrative review of the literature found that AI is less accurate as a diagnostic tool for people with darker skin (Fitzpatrick skin types IV through VI) compared to lighter skin (types I through III), with poorer precision and sensitivity scores across multiple evaluation methods.2PubMed Central. Exploring the Diagnostic Capability of Artificial Intelligence in Dermatology for Darker Skin Tones: A Narrative Review

Specific numbers make the gap concrete. Evaluations using a diverse image dataset found that popular diagnostic models scored meaningfully lower on darker skin. One model’s accuracy metric dropped from 0.64 on the lightest skin tones to 0.55 on the darkest. Another fell from 0.72 to 0.57. Part of the problem is the data itself: some commonly used training databases contain as little as 4.3% images of Black or African American skin.3Journal of Investigative Dermatology. AI Bias in Healthcare Examples: Impacts and Ethical Issues Even when researchers try to fix this by sampling proportionally across skin types, classifiers still underperform for darker skin, suggesting the issue runs deeper than simple representation counts.4arXiv. Predictive Representativity: Uncovering Racial Bias in AI-based Skin Cancer Detection

This matters for real patients. Melanoma is already diagnosed at later stages in people with darker skin, partly because clinical training historically focused on how lesions appear on lighter skin. An AI tool that repeats this blind spot does not just fail to help: it risks reinforcing the very disparity it was supposed to address.

Bias Embedded in Clinical Notes

AI does not only learn from structured data like lab values and billing codes. Increasingly, natural language processing models are trained on the free-text notes that clinicians write in electronic health records. Those notes carry their own biases. A study analyzing ICU notes found that racial and ethnic group descriptors carry different contextual relationships to stigmatizing language, and that these associations shift depending on when and where the notes were written. The researchers concluded that NLP models can transmit implicit bias from their training data, meaning that clinical prediction tools built on those models could reinforce disparities rather than reduce them.5PubMed Central. Measuring Implicit Bias in ICU Notes Using Word-Embedding Neural Network Models

A separate study using Word2Vec embeddings, a type of language model, found troubling associations when probed with analogy questions related to mental health. When the model was asked to complete the pattern “White is to Depression, as Black is to ___,” it returned “undergone electroshock therapy.” When the pattern was “Man is to Depression, as Woman is to ___,” it returned “perinatal depression.” The model was reflecting patterns absorbed from large text corpora, not expressing medical knowledge, but these are exactly the kinds of associations that can shape automated triage or screening if the models are deployed without scrutiny.6PLOS ONE. Artificial Intelligence in mental health and the biases of language based models

Genomic Risk Scores and Ancestry

Polygenic risk scores combine the effects of many genetic variants to estimate someone’s risk for diseases like heart disease, diabetes, or certain cancers. They are a centerpiece of the precision-medicine movement. But the genome-wide association studies used to build these scores have overwhelmingly enrolled participants of European descent. The consequence is that polygenic risk scores are several times more accurate for people of European ancestry than for people of other ancestries.7PubMed Central. Clinical use of current polygenic risk scores may exacerbate health disparities

What makes this different from, say, a drug that happens to work better in one population is the scope. Individual medications might vary in effectiveness across groups, but polygenic risk scores systematically perform better for European-descent populations across nearly every condition. If these scores are rolled into clinical decisions about who gets early screening or preventive treatment, the populations that benefit least from the tool are the ones who often already face the greatest barriers to care. The bias here is not in the AI algorithm per se but in the genomic databases the algorithm depends on.

Hardware Bias That Feeds Into AI

Sometimes the bias enters the pipeline before any algorithm runs at all. Pulse oximeters, the fingertip devices that estimate blood oxygen levels, have long been known to overestimate oxygen saturation in people with darker skin. This measurement error has been documented for decades, yet global use of pulse oximetry continues to increase, and oximeter readings get incorporated into risk scores, triage algorithms, and AI systems.8JAMA. Addressing Racial and Ethnic Bias in Pulse Oximeters—A Wicked Problem

An AI model trained on data from pulse oximeters inherits this hardware-level error. If the model learns that patients with oxygen readings above a certain threshold are “stable,” it may systematically underestimate the severity of hypoxemia in darker-skinned patients. During the COVID-19 pandemic, this became a life-or-death concern as oxygen levels guided treatment decisions. The problem illustrates how bias can be layered: imperfect hardware feeds imperfect data into models that then make imperfect predictions, with each layer compounding the original error.

Age Bias in Medical Imaging

Most public medical imaging datasets used to train AI models contain images from adults. Children are severely underrepresented. When researchers trained cardiomegaly classifiers (models that detect an enlarged heart) on four adult chest X-ray datasets and then tested them on healthy children, the models showed a strong age-related bias. For patients under 11 years old, the rate of false positives increased sharply for younger children. The correlation was strong and statistically clear, and it appeared regardless of which adult dataset the model was trained on.9PubMed Central. Lack of children in public medical imaging data points to growing age bias in biomedical AI

In practical terms, this means an AI reading a chest X-ray of a five-year-old might flag the image as showing an enlarged heart when the heart is perfectly normal for the child’s body size. Pediatric anatomy differs from adult anatomy in ways that adult-trained models simply have not learned. The risk is not just wasted follow-up tests but unnecessary anxiety for families and, in busy clinical settings, alert fatigue that causes real abnormalities to be overlooked.

Gender and Cardiovascular Risk Assessment

Large language models are being explored for clinical decision support, and they bring their own patterns. A study testing GPT-4 on hypothetical cardiovascular risk scenarios found that the model’s assessments shifted based on gender and the presence of psychiatric conditions. In five scenarios without psychiatric comorbidities, GPT-4 rated women as having a higher risk of obstructive coronary artery disease in every single case, pointing to women’s higher age as the decisive factor. But when psychiatric conditions were added to the scenarios, the model’s reasoning flipped: it then rated men as higher risk in more than half the cases.10PubMed Central. Gender Bias in AI’s Perception of Cardiovascular Risk

The concern is not that GPT-4 is being used to diagnose patients today, but that clinicians and health systems are actively exploring such uses. If a model’s risk assessment swings dramatically based on whether a patient has a psychiatric diagnosis, and if that swing differs by gender, the model could steer clinical attention in ways that do not track with actual disease risk. Cardiovascular disease in women is already underdiagnosed relative to men, so a tool that miscalibrates risk by gender could worsen existing gaps.

Models That Do Not Travel Well

AI models trained in wealthy countries are increasingly being deployed in low- and middle-income countries, where healthcare infrastructure, patient populations, and disease patterns differ substantially. A global review found that the vast majority of AI datasets used in published research came from high-income countries, with the United States alone contributing about 41% of all datasets, followed by China at roughly 14%.11PLOS Digital Health. Sources of bias in artificial intelligence that perpetuate healthcare disparities—A global review

When researchers tested whether models developed in the United Kingdom could perform adequately in hospitals in Vietnam, they found clear performance drops. One standard neural network model achieved scores ranging from 0.853 to 0.901 across UK sites but dropped to 0.723 at the Vietnamese hospital. A different model type (XGBoost) held up better, achieving 0.836, but still fell below its UK range.12Scientific Reports. Mitigating machine learning bias between high income and low–middle income countries for enhanced model fairness and generalizability The broader study framing made the point clearly: the effectiveness of AI systems can be compromised when the unique local contexts and requirements of low- and middle-income settings are not adequately considered.13PubMed Central. Generalizability assessment of AI models across hospitals in a low-middle and high income country

This is not just an academic concern. Health systems in lower-resource settings are the ones most likely to adopt off-the-shelf AI tools because they lack the data infrastructure and funding to build their own. If those tools were validated on populations that look nothing like the patients they will serve, the gap between promise and performance can be wide enough to cause harm.

Language Barriers Amplified by AI

AI-powered translation tools are being adopted for medical interpreting, but current systems introduce serious risks for speakers of less common languages. Automatic translation models often struggle with regional varieties, figurative speech, culturally embedded meanings, and emotionally sensitive conversations about topics like reproductive health or chronic disease. These failures can lead to clinically significant misunderstandings.14PubMed Central. Ethical Risks and Structural Implications of AI-Mediated Medical Interpreting

The quality gap between languages is stark. Translation models perform well for widely spoken languages with large digital footprints but degrade for so-called low-resource languages spoken by millions of people but underrepresented in training data. A patient describing symptoms using an idiomatic expression in Hmong or Somali may receive a translation that is nonsensical or, worse, medically misleading. In sensitive clinical contexts, a mistranslation about medication dosage or surgical consent is not just inconvenient. It is dangerous.

When Clinicians Trust the Machine Too Much

Bias in AI tools interacts with a well-documented human tendency called automation bias: the inclination to defer to a computer’s recommendation even when it conflicts with your own judgment. Research on clinical decision support systems has found that when an AI provides an incorrect alert, clinicians’ interpretation accuracy can drop dramatically. In one set of findings, incorrect AI recommendations led to accuracy declines of roughly 43% among one group and 59% among another. Participants showed a 57% increase in errors of commission (acting on a wrong recommendation) when alerted by an incorrect system, compared to working without the system at all.15ScienceDirect. Exploring the risks of automation bias in healthcare artificial intelligence applications: A Bowtie analysis

This creates a compounding problem. If an AI tool carries racial or gender bias, and clinicians tend to follow the tool’s recommendations without sufficient pushback, the bias gets enacted in patient care with the clinician’s implicit endorsement. An analysis of the ethical landscape highlighted that black-box AI systems raise acute concerns around transparency, informed consent, and the risk that automation bias and professional de-skilling could interact to amplify harm rather than catch it.16PubMed Central. Ethical and Legal Challenges of Partially and Fully Autonomous AI in Healthcare: Reinterpreting Liability and Preserving Trust

The Accuracy-Fairness Trade-Off

One reason bias is hard to fix is that efforts to make AI fairer can come at a measurable cost to overall accuracy. When researchers benchmark fairness-correction techniques in real clinical settings, they have found that optimizing for equal performance across racial or gender groups can result in worse model performance overall or in suboptimal calibration for specific subgroups.17PubMed Central. Algorithm fairness in artificial intelligence for medicine and healthcare A separate empirical study found that procedures designed to equalize prediction distributions across groups induced nearly universal degradation of multiple performance metrics within those groups.18PubMed Central. An empirical characterization of fair machine learning for clinical risk prediction

This is a genuinely difficult problem, not just a technical one. If making a model fairer for one group means it performs worse for everyone, hospitals and regulators face uncomfortable questions about whose outcomes to prioritize. There is no universally agreed-upon definition of “fairness” in this context, either: equal accuracy across groups, equal rates of false positives, equal rates of false negatives, and equal positive predictive values are all distinct goals, and satisfying all of them simultaneously is mathematically impossible in most real-world scenarios. The trade-off does not mean fairness efforts are futile, but it does mean that “just debias the algorithm” is not as simple as it sounds.

Regulatory Gaps and the Safety-Net Problem

The U.S. Food and Drug Administration has cleared hundreds of AI-enabled medical devices, but the regulatory framework has struggled to keep pace with the unique risks these tools pose. AI systems that continue learning after deployment can evolve beyond their initial validation, potentially degrading in performance or developing new biases. Although the FDA has proposed regulatory adjustments, significant weaknesses remain in real-time monitoring, transparency requirements, and bias mitigation.19PubMed Central. The illusion of safety: A report to the FDA on AI healthcare product approvals

The institutions most vulnerable to these gaps are the ones serving the most marginalized patients. Safety-net healthcare systems in the United States, which disproportionately serve low-income and historically underrepresented communities, face the highest barriers to evaluating and mitigating algorithmic bias. They have less funding, less technical staffing, and less infrastructure than wealthier hospital systems. The communities they serve are often the very populations underrepresented in the commercial models being deployed. Without concrete, accessible pathways to identify and correct bias, safety-net patients risk falling further behind as AI adoption accelerates.20Nature. Identifying and mitigating algorithmic bias in the safety net

The gap creates a perverse dynamic. The hospitals with the greatest need for efficient, automated tools are also the ones least equipped to catch when those tools go wrong. And the patients who stand to be harmed most are the ones with the fewest alternatives if the system fails them.

Leave a Reply

Your email address will not be published. Required fields are marked *