What Is a Good Number Needed to Treat (NNT)?

A “good” number needed to treat has no fixed threshold because the number only means something in context. An NNT of 20 for a drug that prevents fatal strokes is excellent; an NNT of 20 for a pill that slightly reduces mild headaches, with real side effects, may not be worth it. What makes an NNT good or bad depends on the severity of the condition being prevented, the harms and costs of the treatment, and the baseline risk of the person being treated. The number on its own tells you surprisingly little, which is both its greatest strength and its most common source of misunderstanding.

How NNT Works and Why It Resists Simple Benchmarks

The NNT tells you how many people need to receive a treatment for one additional person to benefit compared to a control group. An NNT of 10 means you treat 10 people and, on average, one of them avoids a bad outcome who would have experienced it without treatment. The other nine either would have been fine anyway or still experience the outcome despite treatment. Lower is better: an NNT of 1 would mean the treatment works for everyone (virtually unheard of in medicine), while an NNT of 100 means you treat a hundred people for one person to benefit.

The reason there is no universal cutoff for “good” is that NNT is tied to the absolute difference in risk between treated and untreated groups. That absolute difference shifts depending on who is being treated. NNT is inversely proportional to the absolute risk reduction, so as the baseline risk of the population changes, the NNT changes too, even when the treatment’s relative effectiveness stays the same.1Europe PMC. Understanding number needed to treat (NNT): A practical guide for anaesthesia and critical care clinicians A blood pressure drug that shows an NNT of 30 in high-risk patients could easily show an NNT of 200 in low-risk patients, because the low-risk group had fewer strokes to prevent in the first place. The drug did not become less effective; the population changed.

Rough Ranges That Clinicians Actually Use

While no official scale exists, experienced clinicians tend to think about NNTs in loose tiers based on clinical stakes. For acute, life-threatening conditions, NNTs in the single digits are common and expected. Giving clot-busting drugs during a heart attack, for example, typically produces NNTs well under 50. For chronic preventive therapies taken over years, NNTs in the range of 20 to 100 are typical and often considered acceptable, particularly when the event being prevented is severe. When you get into the hundreds or thousands, you are usually looking at screening programs or population-level interventions where the disease is relatively rare but the screening is cheap and safe.

Statin therapy for heart disease is one of the most studied examples and nicely illustrates how range and risk interact. In people who already have cardiovascular disease, statins significantly reduce deaths and major coronary events.2PubMed. Comparative benefits of statins in the primary and secondary prevention of major coronary events and all-cause mortality: a network meta-analysis of placebo-controlled and active-comparator trials But in primary prevention, where healthy people take statins to avoid a first event, the NNTs vary enormously depending on genetic and clinical risk. An analysis of two large statin trials found that the ten-year NNT to prevent one coronary heart disease event ranged from around 20 in people at high genetic risk to around 60 in those at low genetic risk.3The Lancet. Genetic risk, coronary heart disease events, and the clinical benefit of statin therapy: an analysis of primary and secondary prevention trials An NNT of 20 to prevent a heart attack over a decade sounds compelling. An NNT of 66 for the same drug in someone at lower risk is harder to get excited about, especially when you factor in years of daily pills and a nonzero chance of side effects like muscle pain.

Screening programs sit at the far end of the spectrum. The number needed to screen to prevent one death from colon cancer with faecal occult blood testing was about 1,374 over five years, and for mammography to prevent one breast cancer death in women aged 50 to 59, it was roughly 2,451 over five years.4PubMed Central. Number needed to screen: development of a statistic for disease screening Those numbers sound enormous, but the screening itself is inexpensive, minimally invasive, and carries low risk. The calculation changes when you are talking about colonoscopies rather than stool tests, because the procedure carries some risk and considerably more cost. In that context, larger NNTs become harder to justify.

The Threshold That Actually Matters

Rather than asking “is this NNT good,” the more useful question is whether the NNT sits below or above the point where the treatment’s benefits outweigh its harms. Researchers have formalized this idea as the “threshold NNT,” the number at which the benefit of treating that many patients exactly equals the negative consequences of treating them. If your treatment’s actual NNT falls below that threshold, the treatment is worth offering. If it falls above it, the harms or costs outweigh the gains.5Journal of Clinical Epidemiology. When should an effective treatment be used?: Derivation of the threshold number needed to treat and the minimum event rate for treatment

This framing clarifies why NNT on its own is incomplete. You also need to know the number needed to harm, or NNH, which works identically but for side effects: how many people must be treated before one additional person experiences a specific adverse event. A treatment with an NNT of 30 and an NNH of 500 looks attractive, because you help a lot more people than you hurt. A treatment with an NNT of 30 and an NNH of 35 is a much tougher call. And if the harm is mild (say, temporary nausea) while the benefit is preventing death, you might accept an NNH that is actually lower than the NNT.6PubMed. When does a difference make a difference? Interpretation of number needed to treat, number needed to harm, and likelihood to be helped or harmed

Severity matters enormously in this calculation. Preventing a death with an NNT of 100 is a bargain most people would take. Preventing a mild rash with an NNT of 5 might not be worth much if the treatment causes fatigue in everyone. The numbers do not exist in a vacuum, and interpreting them requires knowing what is at stake on both sides of the equation.

How Doctors and Patients Misread the Number

One of the most persistent misunderstandings about NNT is the belief that it means only one patient out of that number actually benefits. In a survey of physicians, about a third would not prescribe a drug based on its NNT, and of those, over three-quarters agreed with the statement that “only one out of NNT patients benefits from the treatment.”7PubMed. Medical doctors’ perception of the “number needed to treat” (NNT). A survey of doctors’ recommendations for two therapies with different NNT That framing is misleading. The NNT describes the average difference between two groups. Every treated patient had a reduced probability of the event; it is not the case that one person was “helped” and the rest received no benefit whatsoever. The drug may have slightly lowered the risk for everyone while only crossing the observable threshold for one person.

Patients, too, struggle with NNT. A randomized trial presented people with NNTs ranging from 50 to 1,600 and asked whether they would consent to therapy. Without any explanation of what the numbers meant, about 70% of participants agreed regardless of the NNT. The trend toward lower consent at higher NNTs was weak and not statistically significant. Once participants received an interpretation of the NNT, overall consent dropped from about 70% to 49%, but the trend across NNT levels flattened out entirely.8JAMA Internal Medicine. Decisions on Drug Therapies by Numbers Needed to Treat: A Randomized Trial In other words, people who understood the number were more cautious across the board, but they did not reliably distinguish between an NNT of 50 and an NNT of 1,600. The statistic simply does not translate intuitively for most people without considerable framing.

Even among physicians, understanding is uneven. A study of GPs and hospital physicians found that a majority could correctly identify that increasing patient age would reduce the NNT for a preventive treatment, since older patients have a higher baseline risk.9Family Practice. GPs’ and physicians’ interpretation of risks, benefits and diagnostic test results But correctly answering one conceptual question is a long way from applying NNTs fluently in clinical conversations, and the broader literature suggests that misinterpretation is common enough to influence prescribing decisions in the real world.

Why Reported NNTs Are Often Unreliable

Even if you know how to interpret an NNT, you may be working with a number that was calculated incorrectly. A review of published articles that derived NNTs from studies with time-to-event outcomes, the kind of data you get when a trial tracks how long patients go before experiencing an event, found that only about half used an appropriate calculation method.10PubMed Central. Calculation of NNTs in RCTs with time-to-event outcomes: a literature review In many cases, researchers simply divided one by the raw difference in event rates at the end of the study, which ignores the fact that patients drop out at different times and events accumulate unevenly. For time-to-event data, more sophisticated survival-analysis methods are needed, and when those methods are skipped, the resulting NNTs can be substantially off.

Meta-analyses add another layer of trouble. When multiple trials of different lengths get pooled into a single NNT, the result can be misleading. A meta-analysis of 22 trials of tiotropium in chronic obstructive pulmonary disease reported a single NNT of 16 “over one year,” but the underlying trials ranged from 3 to 48 months, with actual NNTs varying from 15 to 250.11PubMed Central. The Number Needed to Treat: 25 Years of Trials and Tribulations in Clinical Research Collapsing that range into a single figure obscures important variation. The same paper noted a trial of elderly hypertensives that reported an NNT of 94 to prevent one stroke over two years, but a more careful calculation yielded an NNT of 63 — a difference that matters for clinical decisions.

There is also the problem of Simpson’s paradox, a statistical quirk where combining data from multiple trials can reverse the apparent direction of a treatment effect. Researchers have shown that simply lumping trial data together as if it came from one big study can produce NNTs that are biased or even point in the wrong direction when there are imbalances between treatment groups across the included trials.12PubMed Central. Meta-analysis, Simpson’s paradox, and the number needed to treat In one published Cochrane review of nursing interventions for smoking cessation, formal meta-analysis showed a benefit, but pooling the raw patient data naively reversed the effect entirely.13PubMed. Simpson’s paradox and calculation of number needed to treat from meta-analysis The correct approach uses pooled estimates from proper meta-analytic methods rather than treating all the participants as if they were in one giant trial.

NNT in Guideline Decisions

Clinical guidelines increasingly try to incorporate NNT-based reasoning, particularly for preventive therapies where millions of people might be treated. A comparison of five major statin guidelines found something counterintuitive: although the guidelines differed markedly in how many people they would recommend statins for, the estimated NNTs to prevent one cardiovascular event over ten years were nearly identical across all five.14JAMA Cardiology. Statin Use in Primary Prevention of Atherosclerotic Cardiovascular Disease According to 5 Major Guidelines for Sensitivity, Specificity, and Number Needed to Treat The more liberal guidelines cast a wider net but still managed to recommend treatment primarily for people at high enough risk that the NNT remained reasonable. This is reassuring in one sense (the guidelines are not recommending treatment to people who would barely benefit), but it also shows why NNT alone cannot resolve disagreements about how aggressively to treat.

Some researchers have argued that the NNT has outlived its usefulness as a decision-making tool entirely. A recent paper in the Journal of Clinical Epidemiology highlighted important limitations when using NNT for guideline development, policy decisions, or health technology assessment, arguing that the measure’s intuitive appeal sometimes obscures the more nuanced information policymakers need.15Journal of Clinical Epidemiology. The number needed to treat: it is time to bow out gracefully The core complaint is that NNT collapses a complex tradeoff into a single integer, and that integer behaves in ways that can trip up even experts. For instance, the confidence interval around an NNT is notoriously hard to interpret: it can span from positive infinity to negative infinity when the absolute risk reduction is close to zero, creating the bizarre situation where you cannot tell whether the treatment helps or harms.

When NNT Gets Paired with Cost

A natural extension of NNT is to multiply it by the cost of treatment, yielding the “cost per event avoided.” If a drug costs $1,000 per patient per year and the NNT over ten years is 50, you are spending roughly $500,000 to prevent one event. Whether that sounds like a lot depends on what the event is. Preventing one death for $500,000 is well within the range that most health systems consider acceptable. Preventing one case of moderate joint pain for $500,000 is not.

Recent work has explored this idea using composite endpoints. For GLP-1 receptor agonists used in cardiometabolic disease, the cost per event avoided dropped substantially when the definition of “event” was broadened to include more outcomes. Using a narrow three-component endpoint, the cost per avoided event for semaglutide was roughly $1.66 million, but expanding to a broader cardiometabolic endpoint brought it down to about $190,000.16PubMed. Calculating cost per event avoided using a composite number needed to treat This makes intuitive sense: a drug that prevents multiple types of bad outcomes simultaneously is a better deal than one that only prevents one, even if the individual NNTs are identical.

But using NNT as the backbone of economic analysis has real problems. Because NNT does not capture when events occur, only whether they occur within the study window, it treats a treatment that delays a heart attack by ten years the same as one that prevents it entirely. A Health Economics paper argued that this limitation makes NNT-based cost-effectiveness analyses unreliable, particularly for treatments whose main benefit is delaying rather than completely preventing adverse events.17PubMed. Cost-effectiveness analysis based on the number-needed-to-treat: common sense or non-sense? The NNT also forces you into a single binary outcome. You cannot easily capture quality-of-life improvements or partial benefits, which is why formal health-economic analyses typically rely on different tools like quality-adjusted life years rather than NNT alone.

Competing Risks and the Real World

In clinical trials, patients sometimes experience an event that prevents the outcome of interest from ever occurring. If a study tracks time to first heart attack, a patient who dies of cancer during the trial never had a chance to have the heart attack the study was looking for. These “competing risks” complicate NNT calculations because standard methods assume that everyone who did not experience the event was still at risk for it, which is not true if some of them died of something else first. Researchers have proposed methods for computing NNT in the presence of competing risks using cumulative incidence functions, which account for the fact that some people exit the risk pool for reasons other than the treatment working.18PubMed. Number needed to treat for time-to-event data with competing risks This matters most in older populations or in patients with multiple serious conditions, exactly the groups where preventive treatments are most frequently prescribed.

The upshot is that an NNT derived from a trial of relatively healthy 55-year-olds may not translate well to a sicker, older group, even though the sicker group has a higher baseline risk. Higher baseline risk lowers the NNT, but competing risks can offset that advantage if the patients are dying of other causes before the treatment has a chance to work. This is one of the reasons geriatricians and oncologists are often cautious about extrapolating NNTs from landmark trials to their own patient populations.

What Visual Communication Research Suggests

Given how poorly raw NNT numbers communicate to patients and even some clinicians, there has been growing interest in visual tools. Icon arrays, where 100 or 1,000 small figures are shown with a handful highlighted to represent those who benefit, tend to help people grasp absolute risk more accurately. Research on patient decision aids suggests that simple visual formats can reduce common judgment biases and improve understanding, though their effectiveness depends on the viewer’s baseline comfort with numbers.19PubMed Central. Current Challenges When Using Numbers in Patient Decision Aids: Advanced Concepts A person with high graph literacy benefits more from a well-designed chart than someone who struggles to read bar graphs in the first place, which means there is no one-size-fits-all presentation format.

Some decision-aid designers have moved away from presenting NNT as a single number and instead show something like: “Out of 100 people like you who take this drug for 10 years, 3 will avoid a heart attack. Out of those same 100, 97 would have been fine either way or will have a heart attack despite the drug.” That kind of natural-frequency framing tends to be understood more easily than saying “the NNT is 33.” It grounds the statistic in something concrete, and it avoids the trap of making the number sound either impressively small or discouragingly large without context.

The broader lesson is that NNT was designed as a communication tool in the first place. It was introduced in the 1980s specifically to help clinicians and patients think about treatment benefits in absolute rather than relative terms, because relative risk reductions sound much more dramatic. A drug that cuts your risk “in half” sounds transformative; the same drug described as moving your ten-year event rate from 4% to 2%, producing an NNT of 50, sounds more modest. Both descriptions are accurate. The NNT was meant to provide the honest, grounded version. Whether it actually succeeds at that depends entirely on whether the people reading it understand what they are looking at.