Diagnostic Reasoning: A Step-by-Step Clinical Example

Diagnostic reasoning is the mental process a clinician uses to move from a patient’s initial complaints to a working diagnosis and a plan of action. It is not a single skill but a chain of cognitive steps: forming early hypotheses within moments of meeting a patient, gathering targeted information to test those hypotheses, updating the probability of each possibility as new data arrives, and deciding when the evidence is strong enough to act on. A running clinical scenario helps make each step concrete, so picture a 58-year-old woman who arrives at a clinic with new-onset chest tightness and shortness of breath over the past two days.

The First Few Minutes and Early Hypotheses

Within the first seconds of a clinical encounter, an experienced clinician’s brain is already sorting possibilities. The patient’s age, sex, appearance, and chief complaint activate stored patterns, sometimes called illness scripts, that link clusters of findings to specific diagnoses. A 58-year-old woman with chest tightness and dyspnea immediately raises a short mental list: is this a heart problem, a blood clot in the lung, a respiratory infection, an anxiety episode, or something else? This early, rapid sorting is largely intuitive, drawing on pattern recognition built through years of seeing similar patients.

Cognitive psychologists describe this as part of a “dual-process” framework. One track is fast, automatic, and pattern-driven. The other is slow, deliberate, and analytical. In real practice, both tracks run simultaneously and interact constantly. A clinician might recognize a pattern instantly but then switch to methodical analysis when something does not fit the expected picture.1Medical Education Online. An analysis of clinical reasoning through a recent and comprehensive approach: the dual-process theory The interplay between intuition and analysis is not a weakness of the process; it is the core design feature that allows clinicians to be both fast and thorough.

Research on how this initial hypothesis generation works found that clinicians typically generate several diagnostic possibilities within the first few minutes of a clinical encounter based on early cues, then gather data to rule each one in or out. This cycle involves picking up clues, generating hypotheses, interpreting new information, and re-evaluating the hypotheses, all running in a loop rather than a straight line.2PubMed Central. Models of clinical reasoning with a focus on general practice: A critical review

Building a Problem Representation

Once a clinician has gathered the initial history, they do something that often goes unnoticed: they mentally compress all the raw details into a concise summary called a problem representation. For our 58-year-old patient, this might sound like: “A postmenopausal woman with a history of hypertension and recent long car travel presenting with two days of progressive exertional dyspnea and pleuritic chest tightness, without fever or cough.” That single sentence does a lot of work. It flags the relevant risk factors (hypertension, prolonged immobility), highlights the key positive findings (exertional dyspnea, pleuritic quality), and quietly notes what is absent (no fever, no cough), which helps steer away from pneumonia.

This skill is not just a communication shortcut for talking to colleagues. It actively shapes the clinician’s own thinking. A well-constructed problem representation guides which diagnoses stay on the list and which get demoted. Coaching frameworks used in medical training emphasize three elements for sharpening this skill: clarity of the synthesized clinical problem, emphasis on diagnostically relevant positive and negative findings, and use of precise medical terminology that maps to known disease patterns.3PubMed Central. How to … Coach Problem Representation to Strengthen Diagnostic Reasoning in Trainees A vague problem representation leads to vague thinking. A sharp one keeps the reasoning on track.

How Predictions Shape the Physical Exam

Here is something that surprises many people: the diagnosis a clinician has in mind before examining you changes how accurately they interpret what they find. In a study of medical students performing cardiac auscultation, those who correctly predicted the diagnosis from the clinical history before listening achieved an accuracy rate of about 87%, compared to roughly 30% for those who predicted the wrong diagnosis beforehand. Students who received no clinical context at all landed in between, at about 54%.4Europe PMC. Influence of predicting the diagnosis from history on the accuracy of physical examination

The takeaway is double-edged. A good hypothesis sharpens the physical exam. A wrong one can actually make performance worse than having no hypothesis at all. For our chest-tightness patient, a clinician who suspects a pulmonary embolism will listen carefully for a right-sided heart strain pattern and check for leg swelling, and they will be more attuned to subtle findings that match. But if the clinician has prematurely locked onto heart failure, those same subtle PE clues could be overlooked or misinterpreted.

Updating Probability With Each New Piece of Data

Every test result, physical finding, and piece of history either raises or lowers the probability of each diagnosis on the mental list. This is Bayesian reasoning at its core, and clinicians use it constantly even when they do not think of it in mathematical terms. Before any testing, the clinician starts with a rough sense of how likely each diagnosis is based on the patient’s presentation, called the pretest probability. A positive or negative test result then shifts that probability up or down.

The amount of shifting depends on how powerful the test is, captured by something called a likelihood ratio. A test with a high positive likelihood ratio can dramatically increase the probability of disease when it comes back positive, which is especially important in situations where the pretest probability is low and a clinician needs strong evidence before committing to a diagnosis.5Current Opinion in Epidemiology and Public Health. Bayesian interpretation of likelihood ratios in clinical diagnosis and public health For our patient, a D-dimer blood test might be ordered. If it comes back negative in a patient with low pretest probability for PE, the diagnosis is essentially ruled out. If she has high pretest probability and the D-dimer is positive, imaging is the next step, because the test alone does not push the probability high enough to start treatment.

Clinicians do not literally calculate these numbers at the bedside for every patient. But the logic is baked into clinical guidelines and decision rules. The Wells score for PE, for instance, is a structured way to estimate pretest probability so that the right test gets ordered for the right patient.

Deciding When to Test, Treat, or Wait

Not every uncertain diagnosis requires more testing. A landmark concept in clinical decision-making describes two threshold probabilities that govern the clinician’s next move. Below a “testing threshold,” the probability of disease is so low that the risks of further testing outweigh the benefits, and treatment is simply withheld. Above a “test-treatment threshold,” the probability is so high that treatment should begin without subjecting the patient to additional diagnostic procedures. Testing is only worthwhile when the probability falls between those two thresholds.6PubMed. The threshold approach to clinical decision making

Where clinicians set those thresholds varies with experience. A study comparing medical students, residents, and faculty found testing thresholds of roughly 16%, 21%, and 26% respectively, while treatment thresholds hovered around 72% to 79%. The gap between the two thresholds shrank with experience, from about 63 percentage points for students down to 48 for faculty, meaning experienced clinicians had a narrower zone of uncertainty in which they felt testing was needed.7PubMed. Dealing with uncertainty in clinical reasoning: A threshold model and the roles of experience and task framing About two-thirds of respondents in that study made decisions consistent with their own thresholds, while roughly a quarter escalated beyond what their thresholds predicted, ordering more tests or starting treatment earlier than the probabilities alone would suggest.

For our patient with chest tightness, if the clinical picture and Wells score place PE probability squarely in the middle zone, testing with CT angiography is the right call. If the probability is extremely low, the clinician might reasonably reassure her and monitor symptoms. The threshold framework gives structure to what might otherwise feel like gut instinct.

Where Reasoning Goes Wrong

Cognitive biases are the most studied failure points in diagnostic reasoning, and anchoring bias gets the most attention for good reason. Anchoring happens when a clinician latches onto an early piece of information and fails to adjust adequately as new data arrives. Research estimates that cognitive biases contribute to up to 75% of diagnostic errors.8PubMed. Impact of anchoring bias in medical diagnostic decision-making: an experimental study

A large study using real emergency department data showed how this plays out with measurable consequences. When a referring provider mentioned congestive heart failure in the transfer note, emergency physicians were about five percentage points less likely to order testing for pulmonary embolism and took over 15 minutes longer to get around to it when they did. That anchoring nudge toward heart failure did not change the actual rate of PE found, meaning some of those patients truly had PE but had their workup delayed because a different diagnosis was planted early.9JAMA Internal Medicine. Evidence for Anchoring Bias During Physician Decision-Making The finding is unsettling because the anchor was not even a confident diagnosis; it was just a mention in a note.

A related and equally dangerous error is premature closure, which means accepting a diagnosis before the evidence fully supports it and stopping the search too soon. It is considered a key trigger of diagnostic errors.10PubMed. Premature diagnostic closure: an avoidable type of error For our patient, premature closure would look like this: the initial ECG shows a nonspecific abnormality, the clinician calls it “probably cardiac,” starts a cardiac workup, and never circles back to consider PE even when the D-dimer comes back elevated. Interestingly, research on premature closure found that the traits protecting against it were not simply years of experience but rather how many diagnoses a clinician kept on their list, how persistently they sought additional information, and how willing they were to change their leading diagnosis when contradictory evidence appeared.11PubMed. Avoiding premature closure and reaching diagnostic accuracy: some key predictive factors A clinician who carries three plausible diagnoses forward, rather than one, is structurally harder to trap in premature closure.

Experience does help with anchoring to a degree. Emergency medicine faculty had significantly lower anchoring error rates than residents in one study comparing the two groups, suggesting that diagnostic pattern libraries built over years provide some insulation against being pulled off course by misleading cues.12PubMed Central. Anchoring Errors in Emergency Medicine Residents and Faculties

Slowing Down on Purpose

If biases are the disease, metacognition is the main treatment. Metacognition in this context means the clinician’s ability to step back from their own thinking and ask: am I sure about this, or am I cutting corners? One practical tool for this is the “diagnostic pause,” a structured prompt that asks the clinician to reconsider the working diagnosis before finalizing it. In a study implementing diagnostic pauses in outpatient settings, about 13% of the alerts led clinicians to report taking new actions they would not have otherwise taken.13BMJ Quality & Safety. Implementation of diagnostic pauses in the ambulatory setting That may sound modest, but in a high-volume practice, one in eight cases catching a reconsideration is a meaningful safety net.

Other debiasing strategies include deliberately listing “what else could this be” before finalizing a diagnosis, using checklists for high-stakes presentations, and discussing cases with a colleague when something feels off. None of these eliminate bias entirely. The goal is to create friction in the right places so that the fast, intuitive track does not run unchecked when the case is complex or atypical.

Interruptions and the Clinical Environment

Clinical settings are noisy, chaotic, and full of interruptions. You might expect that constant disruption would erode diagnostic accuracy, but the evidence is more nuanced. In a controlled experiment, interruptions did not reduce diagnostic accuracy for either residents or experienced emergency physicians. However, interrupted cases took meaningfully longer to diagnose, with an average of 71 seconds per case compared to 63 seconds without interruption. Experience level, rather than environmental disruption, was the dominant factor in accuracy: emergency physicians were correct about 71% of the time compared to 43% for residents, regardless of whether interruptions occurred.14Academic Medicine. Disrupting Diagnostic Reasoning: Do Interruptions, Instructions, and Experience Affect the Diagnostic Accuracy and Response Time of Residents and Emergency Physicians?

The practical implication is that seasoned clinicians seem to have developed some resilience to environmental chaos, at least for the types of cases studied. But the added time cost is real and cumulative. In a busy emergency department seeing dozens of patients per shift, adding eight extra seconds per case to account for interruption-induced delays adds up. The deeper concern is that real-world interruptions are more varied and sustained than those in controlled experiments, and their effects on complex, multi-step reasoning may not be fully captured.

AI as a Diagnostic Partner

Artificial intelligence tools are increasingly entering the diagnostic process, particularly as clinical decision support systems that suggest diagnoses or flag abnormalities. The evidence so far is genuinely mixed. AI-powered decision support shows promise in processing large amounts of health data, and there are proof-of-concept successes in predicting outcomes from electronic health records or identifying hidden conditions from routine tests. But evidence for consistent improvements in day-to-day clinical practice remains uneven, with human factors like how clinicians interact with and trust AI playing a major role.15PubMed Central. Artificial Intelligence in Clinical Medicine: Challenges Across Diagnostic Imaging, Clinical Decision Support, Surgery, Pathology, and Drug Discovery

One particularly revealing study found that AI recommendations had a strong bidirectional effect on clinician accuracy. Participants were ten times more likely to reach the correct diagnosis when the AI suggestion was correct but became less accurate when the AI suggestion was wrong. The overall effect of introducing AI was essentially a wash compared to baseline performance, because the gains from correct AI advice were offset by the losses from incorrect advice.16International Journal of Medical Informatics. Impact of AI recommendation correctness on diagnostic accuracy in clinical decision-making Clinicians with stronger baseline diagnostic skills were better at resisting incorrect AI nudges, suggesting that AI works best as a tool for already-competent reasoners rather than as a crutch for weaker ones.

Teaching the Reasoning Process

Diagnostic reasoning is often taught implicitly through clinical rotations, but structured methods exist to make the process more visible and teachable. One well-studied approach is SNAPPS, an acronym that guides trainees through summarizing the case, narrowing the differential, analyzing the possibilities, probing the preceptor with questions, planning management, and selecting a learning issue. A systematic review with meta-analysis found that SNAPPS increased the number of differential diagnoses trainees generated and encouraged them to express their uncertainties openly, both of which are protective against premature closure. The benefits appeared more pronounced for residents than for medical students.17PubMed. Effects of SNAPPS in clinical reasoning teaching: a systematic review with meta-analysis of randomized controlled trials

When compared to another common teaching method, SNAPPS was particularly effective at helping students take an active role in case presentation rather than passively reciting facts to an attending physician.18PubMed Central. Case presentation methods: a randomized controlled trial of the one-minute preceptor versus SNAPPS in a controlled setting Making each reasoning step explicit, rather than leaving it invisible inside the trainee’s head, gives preceptors a window into where the thinking is going right or wrong. This matters because the failure points in diagnostic reasoning are usually invisible unless someone deliberately surfaces them.

Training based on illness script theory, which organizes medical knowledge around typical patient stories rather than isolated facts, has also shown measurable effects. An experimental study found that teaching interns to reason through illness scripts significantly improved their clinical reasoning scores and their ability to identify key diagnostic features in patient presentations.19PubMed Central. How to develop clinical reasoning in medical students and interns based on illness script theory: An experimental study

Communicating Uncertainty to Patients

Diagnostic reasoning does not happen in a vacuum. The patient sitting across from you needs to understand what is happening and what to watch for, especially when the diagnosis is still uncertain. Safety-netting is the practice of telling a patient: here is what I think is going on, here is what to watch for, and here is when to come back. But how much uncertainty to share is genuinely tricky. In a study exploring clinician attitudes, some participants worried that openly communicating diagnostic uncertainty during safety-netting could cause patient anxiety or drive unnecessary return visits and over-investigation.20PubMed Central. Role of communicating diagnostic uncertainty in the safety-netting process: insights from a vignette study

There is no clean answer to how much to say. Too little transparency and the patient may not return when warning signs appear. Too much and you risk creating worry without giving the patient tools to manage it. The emerging consensus leans toward structured honesty: tell the patient what the leading possibility is, name the one or two serious alternatives still on the radar, and give specific instructions about what would warrant coming back. “If you develop sudden worsening of your breathing, calf pain, or you cough up blood, go to the emergency room” is more useful than “come back if things get worse.”

When Tests Create New Problems

An underappreciated wrinkle in diagnostic reasoning is the problem of incidental findings. Every imaging study a clinician orders has a chance of finding something unrelated to the original question. A systematic review of diagnostic imaging studies found that incidental findings turned up in an average of about 24% of cases, with the rate climbing to roughly 31% when CT scans were involved.21Europe PMC. Incidental findings in imaging diagnostic tests: a systematic review Of those incidental findings, only about 46% were ultimately confirmed to be clinically relevant after follow-up.

For our 58-year-old patient, a CT angiogram ordered to look for pulmonary embolism might turn up a small thyroid nodule or a lung nodule that has nothing to do with her chest tightness. Now the clinician faces a new reasoning chain: is this nodule something to worry about? Does it need follow-up imaging in six months? Does the patient need to be told in a way that does not generate unnecessary panic? Incidental findings are a byproduct of powerful diagnostic tools, and managing them has become a reasoning challenge in its own right.

Reasoning as a Team Sport

Diagnostic reasoning is often framed as something happening inside one clinician’s head, but in real practice it is frequently collaborative. A nurse notices a subtle change in a patient’s mental status. A pharmacist flags a drug interaction that could explain new symptoms. A radiologist adds a note suggesting an alternative interpretation of an image. Collaborative clinical reasoning allows cognitive load to be shared across multiple professionals, with each person contributing their perspective and catching potential blind spots. This kind of distributed thinking may help identify biases that a solo clinician might miss.22medRxiv. Collaborative Clinical Reasoning: a scoping review

For our patient, the emergency physician might lean toward a cardiac workup based on the ECG, while the nurse who took the history noticed the patient mentioned a recent 12-hour car ride, raising the PE concern. That kind of cross-pollination does not happen automatically. It requires a team culture where all members feel comfortable voicing observations, and where the diagnostic conversation is genuinely open rather than a one-way flow from physician to everyone else. The formal study of team-based reasoning is still relatively young, but the intuition behind it is old: two sets of eyes catch more than one.