What Does Medical Coding Do? How It Powers Healthcare

Medical coding translates every diagnosis, procedure, and clinical encounter into standardized alphanumeric codes that the rest of the healthcare system depends on. These codes determine how hospitals and physicians get paid, feed the databases that track disease outbreaks, shape insurance premiums, and supply the raw data behind most clinical research conducted on large populations. Without coding, a doctor’s note about a patient’s broken wrist or a complicated surgery would remain trapped in free text, invisible to the billing department, the public health agency, and the researcher trying to figure out whether a treatment actually works.

From Clinical Notes to a Universal Language

At its core, medical coding is a translation task. A coder reads through a patient’s medical record and assigns codes from standardized classification systems. The most widely used system for diagnoses is the International Classification of Diseases, currently in its tenth revision in most countries (ICD-10), which contains tens of thousands of codes covering everything from a sprained ankle to a rare autoimmune disorder. In a comparative analysis of clinical decision-support systems, ICD-10 was the coding system used most frequently, appearing in over 93% of diagnosis-related code sets examined.1PubMed Central. Coding Systems for Clinical Decision Support: Theoretical and Real-World Comparative Analysis Procedures have their own code sets, and other specialized terminologies handle drugs, lab results, and medical devices. The coded data then drive hospital reimbursement, audit, research, and benchmarking.2Annals of Surgery. A Study of Clinical Coding Accuracy in Surgery: Implications for the Use of Administrative Big Data for Outcomes Management

The coder’s job sits at a peculiar intersection. They need enough clinical knowledge to understand what the physician documented and enough familiarity with coding rules to pick the right code from thousands of options. A surgeon might write “repaired abdominal aortic aneurysm using endovascular graft” in an operative note; the coder needs to know which specific procedure code matches that description and which diagnosis codes capture the aneurysm, any complications, and any relevant comorbidities. Getting it wrong ripples outward in ways most patients never see.

How Coding Drives Revenue

The most immediate impact of medical coding is financial. The revenue cycle in healthcare tracks the entire payment process from when a patient schedules an appointment through treatment, coding, billing, and eventual reimbursement.3PubMed Central. Revenue Cycle Management: The Art and the Science Codes are the linchpin of that cycle. When a hospital submits a claim to an insurance company or Medicare, the payment amount is calculated based on the diagnosis and procedure codes assigned. A more severe diagnosis or a more complex procedure generally triggers a higher reimbursement, which is why the specificity and accuracy of codes matter enormously.

In the United States, hospitals paid under Medicare’s prospective payment system are reimbursed according to diagnosis-related groups (DRGs), which bundle codes into categories with preset payment amounts. If a coder misses a complication or comorbidity that would bump a case into a higher-paying DRG, the hospital absorbs the loss. One study of patients undergoing elective aortic aneurysm repair found a miscoding rate of about 10%, translating to nearly $588,000 in lost billing opportunity across just 104 patients, with an average loss of roughly $5,400 per correctly coded patient.4PubMed. Financial implications of coding inaccuracies in patients undergoing elective endovascular abdominal aortic aneurysm repair Another audit of 752 cases at a different institution found that about 16% reflected a DRG change after review, adding up to nearly AU$575,300 in additional revenue that had been left on the table.5PubMed. The risk and consequences of clinical miscoding due to inadequate medical documentation: a case study of the impact on health services funding

These aren’t rounding errors. They represent real money that affects a hospital’s ability to hire staff, upgrade equipment, and keep its doors open. The financial stakes explain why hospitals invest heavily in coding departments, auditing programs, and clinical documentation improvement initiatives.

When Documentation Falls Short

Coding accuracy depends almost entirely on the quality of what the physician writes down. If a surgeon performs a complex procedure but documents it vaguely, the coder can only code what’s on paper. The audit that found 16% of cases had DRG changes attributed more than half of those changes to documentation problems rather than coder error.5PubMed. The risk and consequences of clinical miscoding due to inadequate medical documentation: a case study of the impact on health services funding Incorrect selection of the principal diagnosis accounted for another 13%, and missing additional diagnosis codes covered the rest. The takeaway is that coding failures are often documentation failures in disguise.

This is where clinical documentation improvement (CDI) programs come in. These programs train physicians and other clinicians to document more precisely, and they employ CDI specialists who review records and query providers when something looks incomplete. At one academic surgery department, implementing a CDI curriculum that included lectures, one-on-one training, case reviews, and queries from nurse CDI staff and hospital coders led to an estimated $4.7 million increase in charges.6PubMed. Implementation of a Clinical Documentation Improvement Curriculum Improves Quality Metrics and Hospital Charges in an Academic Surgery Department That figure didn’t come from billing for services that weren’t performed. It came from accurately capturing the complexity of services that were already being delivered but weren’t being reflected in the codes.

Upcoding and the Compliance Tightrope

If undercoding costs hospitals money, the reverse problem costs taxpayers money. Upcoding means assigning codes that make a patient’s condition or a provider’s service look more severe or complex than it actually was, resulting in higher reimbursement. The line between aggressive but legitimate coding and outright fraud can be thin, and regulators take it seriously.

The scale of the problem varies across different parts of Medicare. Upcoding in traditional Medicare hospital stays (Part A) has been estimated at about $656 million per year, representing roughly half a percent of total Part A spending. Upcoding in physician services (Part B) runs higher at around $2.4 billion annually, or about 2.4% of Part B expenditures. The biggest concern involves Medicare Advantage plans (Part C), where multiple analyses have pegged upcoding at $10 to $15 billion annually, somewhere between 3% and 4% of Part C spending.7PubMed Central. Upcoding in medicare: where does it matter most? The difference is partly structural: Medicare Advantage plans are paid based on the predicted health costs of their enrollees, so making enrollees look sicker on paper directly increases payments.

Physicians who bill improperly face a range of consequences under programs like the Medicare Integrity Program, including loss of revenue, fraud investigations, financial sanctions, disciplinary action, and exclusion from government programs entirely.8PubMed Central. Addressing medical coding and billing part II: a strategy for achieving compliance. A risk management approach for reducing coding and billing errors The compliance infrastructure around coding, including internal audits, external audits, and whistleblower protections, exists because the financial incentives to push codes upward are real and persistent.

Tracking Disease at a Population Level

Beyond billing, coded data are one of the primary tools public health agencies use to monitor what’s happening across populations. When researchers want to know how many people are dying of heart disease, how common diabetes hospitalizations are in a given region, or whether a new infectious disease is spreading, they typically turn to administrative databases built from coded medical records.

The transition from ICD-9 to ICD-10 in the United States (which happened in 2015 after years of delay) was driven in part by the need for more granular public health data. A study comparing the two systems across reportable diseases, leading causes of death, and terrorism-related conditions found that ICD-10 captured public health diseases significantly more completely and specifically than its predecessor.9PubMed Central. The effectiveness of ICD-10-CM in capturing public health diseases Where ICD-9 might lump several related conditions under one code, ICD-10 often distinguishes them, giving epidemiologists and policymakers more precise information to work with.

That said, the accuracy of coded data for disease surveillance is far from perfect. Codes are assigned primarily for billing, not for epidemiology, and the two goals don’t always align. A study examining how well billing codes captured anaphylaxis found that anaphylactic shock-specific ICD codes had a positive predictive value of only about 52% to 53% when compared against physician chart review. Combining those codes with additional symptom-specific codes improved the figure to roughly 63% to 67%.10PubMed. Capturing anaphylaxis through medical records: Are ICD and CPT codes sufficient? In other words, roughly a third to half of the records flagged as anaphylaxis by billing codes alone didn’t actually represent confirmed anaphylaxis on closer review. Researchers who use administrative data routinely grapple with this kind of noise, developing algorithms that combine multiple codes and clinical criteria to improve accuracy.

Risk Adjustment and Why Your Health Score Matters

Medical codes don’t just describe what happened during a single visit. They’re also used to build a picture of how sick a patient is expected to be in the future. The Centers for Medicare and Medicaid Services uses a risk adjustment model called hierarchical condition category (HCC) coding to estimate the cost and risk associated with providing healthcare to individual patients.11PubMed Central. Is the Centers for Medicare and Medicaid Services Hierarchical Condition Category Risk Adjustment Model Satisfactory for Quantifying Risk After Spine Surgery? Each patient’s diagnostic codes feed into a risk score, and that score determines how much money a health plan or provider receives to care for them.

This is the mechanism behind the Medicare Advantage upcoding concern mentioned earlier. If a plan’s enrollees have higher risk scores, the plan gets higher payments. The system was designed to ensure that plans enrolling sicker patients receive adequate funding, but it also creates a financial incentive to document every possible diagnosis. An annual wellness visit where the doctor carefully reviews and codes a patient’s full list of chronic conditions isn’t just good medicine; it’s worth real money to the plan.

Risk adjustment also plays a growing role in value-based care arrangements, where providers are evaluated and sometimes paid based on patient outcomes. If a surgeon’s complication rate is being compared against a national benchmark, the severity of the patients they treat matters. Accurate coding ensures that a surgeon operating on high-risk patients isn’t unfairly penalized when compared to one treating healthier people. Under-code the complexity of your patient population, and you look like you’re performing worse than you actually are.

Coding Social Determinants of Health

One of the newer frontiers in medical coding involves capturing factors that aren’t strictly medical but profoundly affect health outcomes: housing instability, unemployment, lack of social support, food insecurity, and similar conditions. ICD-10 includes a set of “Z-codes” specifically for these social determinants of health, and their use has been growing, though from a very low baseline.

An analysis of over 14 million hospital admissions in 2016 and 2017 found that only about 1.9% had a social determinant Z-code attached. The most common categories covered problems related to housing and economic circumstances, upbringing, family circumstances, and employment. The proportion of hospitalizations with at least one such code did increase between the two years, but the overall rate remained low.12PubMed Central. Utilization of Social Determinants of Health ICD-10 Z-Codes Among Hospitalized Patients in the United States, 2016-2017

The gap between how common social risk factors are and how often they’re coded is enormous. Housing instability, food insecurity, and financial stress affect a far larger share of hospitalized patients than 2%. The underuse likely reflects several barriers: clinicians may not screen for these issues routinely, there’s no direct reimbursement incentive for documenting them, and the workflow for capturing them hasn’t been standardized. Advocates argue that better capture of social determinants through coding could help hospitals identify patients who need extra support, allocate community health resources more effectively, and eventually factor social risk into payment models so that safety-net providers aren’t disadvantaged for serving harder-to-treat populations.

Medical Coding in Clinical Trials

Outside of hospital billing and insurance, coding plays a quieter but essential role in drug development and clinical research. Clinical trial data, including adverse events reported by participants, medical histories, and concomitant medications, need to be coded using standardized medical dictionaries so that they can be aggregated and analyzed across trial sites and countries. The most commonly used dictionaries for this purpose are MedDRA (for adverse events and medical conditions) and WHO-DDE (for drugs).13PubMed Central. Medical coding in clinical trials

Without this coding step, a headache reported by one site as “cephalalgia” and by another as “head pain” would appear as two different events in the database. Standardized coding groups them under a single term, making it possible to detect safety signals that might otherwise be missed in the noise of free-text reporting. This is one of the ways that coding directly affects drug safety: if adverse events aren’t coded consistently, a dangerous pattern might not emerge until more people are harmed.

Automation and Artificial Intelligence

Given the complexity, volume, and financial stakes of medical coding, the appeal of automating it is obvious. Researchers have been working on natural language processing (NLP) and deep learning systems that can read clinical text and suggest or assign codes automatically. Progress has been real but uneven.

The challenge is the sheer number of codes and the subtlety of clinical language. A review of the field noted that the best deep learning models tested on the full set of roughly 8,900 ICD-9 codes in a major clinical database achieved a best score under 60% on a standard accuracy metric, well below what would be needed for reliable autonomous coding.14npj Digital Medicine. Automated clinical coding: what, why, and where we are? Models perform better when distinguishing broadly between simple and complex cases. One NLP system designed to predict how difficult a case would be to code achieved about 71% accuracy at sorting cases into simple versus complex categories.15PubMed Central. An End-to-End Natural Language Processing Application for Prediction of Medical Case Coding Complexity: Algorithm Development and Validation

The current consensus is that computer-assisted coding (CAC) works best as a tool that supports human coders rather than replacing them. A review of the literature on CAC systems found that they demonstrated value in improving coding accuracy and quality, catching things that might be missed during manual review. Rather than threatening coding professionals, the technology appears to be shifting their role toward reviewing and editing machine-generated suggestions, a job that still requires deep clinical and coding knowledge.16PubMed Central. Computer-assisted clinical coding: A narrative review of the literature on its benefits, limitations, implementation and impact on clinical coding professionals The coder of the near future may look less like a translator and more like an editor, but the expertise remains essential.

ICD-11 and What Comes Next

The World Health Organization released ICD-11 in 2019, and countries are at various stages of planning or beginning their transitions to it. The new edition is the first to be designed for the digital age from the ground up. Its architecture includes a semantic knowledge base called the Foundation, an online coding tool that replaces the old printed index, and an application programming interface that allows health IT systems to access ICD-11 content and services remotely. It also has built-in support for multiple languages and an enhanced ability to capture clinically relevant characteristics of cases in combination.17PubMed Central. ICD-11: an international classification of diseases for the twenty-first century

For the people who actually do coding, the transition will mean another massive learning curve, similar to the ICD-9 to ICD-10 shift that consumed U.S. healthcare for years. But the design improvements in ICD-11 reflect lessons learned from the limitations of earlier versions. The ability to combine codes in structured ways, for instance, could reduce the need for the unwieldy combination codes that proliferated in ICD-10 and improve the precision of coded data for both billing and research purposes.

Whether ICD-11 will also improve the capture of social determinants, support better integration with AI-assisted coding tools, or reduce the gap between billing-driven coding and clinically accurate coding remains to be seen. The infrastructure is more flexible, but the human incentives, documentation habits, and workflow pressures that shape coding accuracy haven’t changed just because the classification system got an upgrade. The technology of coding keeps advancing, and the fundamental tension between coding as a financial act and coding as a clinical record remains as alive as ever.