How to Calculate Cumulative Incidence

Cumulative incidence is calculated by dividing the number of new cases of an event by the total number of people at risk at the start of the observation period. The result is a proportion between 0 and 1, representing the probability that a person in that group will experience the event over the specified time window.1Clinical Epidemiology. An Introduction to Competing Risks in Epidemiology That much is simple enough for a back-of-the-napkin estimate. But the moment people drop out of your study, die from something else, or enter the population at different times, the calculation gets more involved.

The Basic Formula and What Goes Into It

In its simplest form, cumulative incidence is just a fraction. The numerator is the count of individuals who developed the condition (or experienced the event) during the observation period. The denominator is the total number of individuals who were at risk when observation began. If 30 out of 500 factory workers developed hearing loss over five years, the cumulative incidence is 30 divided by 500, or 0.06. You can express that as 6%.

A few details matter here. First, the numerator has to be contained within the denominator. Everyone counted as a new case must have been part of the at-risk group at the outset.2PubMed Central. Incidence rates in dynamic populations Second, “at risk” means genuinely capable of developing the condition. If you’re calculating the cumulative incidence of a first heart attack, people who already had one before the study started don’t belong in the denominator. Third, the result is always a dimensionless proportion somewhere between 0 and 1. It cannot exceed 1, because you can’t have more cases than people at risk.2PubMed Central. Incidence rates in dynamic populations This last point sounds obvious, but it distinguishes cumulative incidence from incidence rate, which uses person-time in the denominator and can technically go above 1.

Why the Time Window Changes Everything

Cumulative incidence is meaningless without a stated time period. Saying “the cumulative incidence of diabetes in this population is 8%” tells you nothing unless you know whether that 8% accumulated over one year, five years, or a lifetime. Longer observation windows almost always produce higher cumulative incidence, simply because people have more time to develop the condition.

A study of ear infections in young children illustrates this clearly. Researchers followed a sample of about 2,500 infants and found the cumulative incidence of a first episode of acute otitis media was roughly 42% by 12 months of age and roughly 71% by 24 months.3International Journal of Pediatric Otorhinolaryngology. The occurrence of acute otitis media in infants. A life-table analysis Same population, same condition, but the two numbers paint very different pictures because the observation window doubled. Whenever you see a cumulative incidence figure, check the time horizon before interpreting it.

This sensitivity to time also means that comparing cumulative incidence across studies only works if those studies observed their populations over similar periods. A five-year cumulative incidence from one cancer registry cannot be placed alongside a ten-year figure from another and treated as an apples-to-apples comparison.

Attack Rates Are Just Cumulative Incidence With a Different Name

During disease outbreaks, you’ll often hear public-health officials report an “attack rate.” Despite the name, this is not a rate in the technical sense. An attack rate is the cumulative incidence of infection or disease in a group observed during an outbreak, calculated by dividing the number of exposed individuals who got sick by the total number of individuals at risk.4PLOS ONE. Attack Rates Assessment of the 2009 Pandemic H1N1 Influenza A in Children and Their Contacts: A Systematic Review and Meta-Analysis If 18 people at a wedding reception ate the potato salad and 12 got food poisoning, the attack rate is 12 out of 18, or about 67%.

The formula is identical to the simple cumulative incidence calculation. The different terminology persists mainly because outbreak investigations have their own tradition and vocabulary. The important thing to understand is that the same logic applies: the denominator is the at-risk group (people who were exposed), the numerator is the subset that got sick, and the result is a proportion.

What Happens When People Drop Out

The simple formula works perfectly when every single person in your group is followed for the entire observation period and you know exactly who got the condition and who didn’t. Real studies rarely deliver that luxury. People move away, withdraw consent, or the study ends before everyone has been observed for the full duration. In epidemiology, these incomplete observations are called censored data.

Once censoring enters the picture, the basic division breaks down. If 50 out of 1,000 people developed the condition but 200 others were lost to follow-up before the observation period ended, simply dividing 50 by 1,000 underestimates the true cumulative incidence, because some of those 200 lost participants might have developed the condition had they stuck around.

The Life Table Approach

One classic way to handle censoring is the life table method, sometimes called the actuarial method. The idea is to break the observation period into intervals, usually equal time segments like months or years. Within each interval, you count how many people were still being observed at the start, how many developed the condition during that interval, and how many were censored. People censored partway through an interval are assumed, on average, to have been at risk for half the interval. You adjust the denominator accordingly and then calculate the probability of developing the condition within each interval. These interval-specific probabilities are then combined across the full observation period to produce the overall cumulative incidence.

The otitis media study mentioned earlier used exactly this approach, applying the life table method to handle children who were lost before reaching their second birthday.3International Journal of Pediatric Otorhinolaryngology. The occurrence of acute otitis media in infants. A life-table analysis It’s a well-established technique that works especially well when event times are grouped rather than recorded to the exact day.

The Kaplan-Meier Approach

When you know the exact time each event occurred, the Kaplan-Meier method is the standard tool. It works similarly to the life table method but recalculates the probability of remaining event-free at every point where someone actually develops the condition, rather than at preset intervals. The result is a step-function survival curve, and the cumulative incidence at any point can be read as 1 minus the survival probability at that point.

Kaplan-Meier estimation is the workhorse of survival analysis and handles ordinary censoring reasonably well. However, it carries a major assumption: it treats censored observations as if they’d behave the same as the people still being followed. When that assumption is wrong, particularly when the reason someone dropped out is related to their risk of developing the condition, the estimate becomes unreliable. This is sometimes called informative censoring, and dealing with it requires sensitivity analyses that incorporate outside information about why people left the study.5PubMed Central. Inference for cumulative incidence functions with informatively coarsened discrete event-time data

Competing Risks Make the Kaplan-Meier Method Unreliable

The Kaplan-Meier approach has a more fundamental problem when people in your study can experience more than one type of event. Imagine you’re studying the cumulative incidence of hip fractures in elderly patients. Some of those patients will die from heart disease or stroke before they ever fracture a hip. Death from another cause is a “competing risk” because it permanently removes the person from the possibility of experiencing the event you care about.

When these competing events exist, the Kaplan-Meier method treats people who die of other causes the same way it treats people who simply drop out of the study. It censors them and carries on as though they might still develop the condition. But they can’t. They’re dead. This leads the Kaplan-Meier estimate to overstate the actual probability of the event occurring, regardless of whether the competing events are statistically independent of the event of interest.6PubMed Central. Introduction to the Analysis of Survival Data in the Presence of Competing Risks

The bias is upward. In other words, using the Kaplan-Meier method in the presence of competing risks will make the condition look more common than it actually is in the real world.6PubMed Central. Introduction to the Analysis of Survival Data in the Presence of Competing Risks The more common the competing event (the more people die of other causes, say), the bigger the overestimate becomes.

The correct approach when competing risks are present is the cumulative incidence function, sometimes abbreviated CIF. Instead of pretending that people who experience a competing event might still develop the condition, the CIF accounts for the fact that they’ve been permanently removed from the at-risk pool by a different outcome. It produces separate cumulative incidence curves for each type of event, and these curves can be stacked. The total across all event types at any given time point represents the overall probability that something happened, which can never exceed 1.7Clinical Cancer Research. Cumulative Incidence in Competing Risks Data and Competing Risks Regression Analysis

The practical takeaway: if your data includes any outcome that prevents people from experiencing the event you’re studying, the Kaplan-Meier complement is the wrong estimator. Use the cumulative incidence function instead. This comes up constantly in geriatric research, oncology, and cardiovascular studies, where death from other causes is always lurking.

Closed Cohorts Versus Dynamic Populations

Cumulative incidence assumes a closed cohort, a group of people who are all identified at a single starting point and followed forward. Nobody new joins partway through. This is what gives the measure its clean interpretation as a probability: everyone shares the same starting line, and the proportion who cross the finish line (develop the event) is the cumulative incidence.2PubMed Central. Incidence rates in dynamic populations

Many real-world populations aren’t closed, though. A hospital’s patient roster changes constantly as people are admitted and discharged. A city’s population shifts with births, deaths, and migration. These are dynamic populations, and they don’t have a single shared starting point. Calculating cumulative incidence in a dynamic population doesn’t work cleanly because the denominator keeps changing.

For dynamic populations, the appropriate measure is the incidence rate (sometimes called person-time incidence rate), which divides the number of new cases by the total person-time contributed by everyone in the population. This measure can handle people entering and leaving at different times because the denominator accumulates each person’s contribution based on how long they were actually observed. The distinction between a closed cohort where cumulative incidence works and a dynamic population where incidence rate is the right tool was already well understood by William Farr in the 1850s, who drew a clear line between what he called the “probability of death” and the “rate of mortality.”

If you do need a cumulative incidence from person-time data, it’s possible to convert under certain assumptions. When event rates are low and roughly constant over time, the cumulative incidence over a period can be approximated from the incidence rate. But this is an approximation, not an exact calculation, and it breaks down when event rates are high or change substantially over the observation period.

Getting Confidence Intervals Around Your Estimate

A single cumulative incidence figure is a point estimate, and like all point estimates, it comes with uncertainty. In a study of 500 people, an observed cumulative incidence of 10% doesn’t mean the true population value is exactly 10%. Confidence intervals give you a range that likely contains the true value.

For the simple formula without censoring, computing a confidence interval is straightforward, using the same methods you’d use for any proportion. Once censoring and competing risks enter the picture, things get more complicated. Standard software can produce confidence intervals around Kaplan-Meier estimates and cumulative incidence functions, but researchers have noted that the methods aren’t always well-calibrated, particularly for the cumulative incidence function. More recent work has developed profile-likelihood and constrained bootstrap approaches that keep the confidence interval within the 0-to-1 bounds that a probability requires.8PubMed. Confidence intervals for the cumulative incidence function via constrained NPMLE

Software programs for these calculations exist in most major statistical platforms. Dedicated routines for computing cumulative incidence with confidence limits, stratified by subgroups and at user-specified time points, are available in packages for SAS, R, Stata, and other environments.9PubMed. A SAS program for calculating cumulative incidence of events (with confidence limits) and number at risk at specified time intervals with partially censored data In R, the commonly used packages include survival (for Kaplan-Meier) and cmprsk or tidycmprsk (for competing-risks cumulative incidence). Review articles have cataloged the available software and regression methods for modeling cumulative incidence functions across platforms.10PubMed Central. Modeling cumulative incidence function for competing risks data

Common Mistakes That Distort the Numbers

Several errors show up repeatedly when people calculate or interpret cumulative incidence, even in published research.

  • Omitting the time window: Reporting “a cumulative incidence of 15%” without specifying the observation period makes the number uninterpretable. Always pair the estimate with its time horizon.
  • Using Kaplan-Meier when competing risks exist: As covered above, this inflates the estimate. The fix is the cumulative incidence function, not the complement of the survival curve.
  • Confusing cumulative incidence with incidence rate: Cumulative incidence is a proportion (cases per people at risk). Incidence rate is a density (cases per person-time). They answer different questions and are not interchangeable. Cumulative incidence tells you the probability of an event over a specific period. Incidence rate tells you the speed at which events occur.
  • Ignoring outcome misclassification: If the condition is measured imperfectly, some true cases will be missed and some non-cases will be incorrectly counted. Both errors distort cumulative incidence. Studies using administrative databases are particularly prone to this, and correction methods exist but are rarely applied in practice.11PubMed. Outcome misclassification: Impact, usual practice in pharmacoepidemiology database studies and an online aid to correct biased estimates of risk ratio or cumulative incidence
  • Applying the formula to a dynamic population: If people enter the at-risk group at different times, the simple division doesn’t give a valid probability. You need a life table method, Kaplan-Meier estimation, or a switch to incidence rates.

When You Actually Need a Different Measure Entirely

Cumulative incidence is the right tool for a specific type of question: what is the probability that a person in this defined group will develop this event over this defined time period? It requires a clear starting point, a defined endpoint, and ideally a closed cohort. When any of those conditions isn’t met cleanly, other measures may serve you better.

If events can recur, like repeated hospitalizations or multiple infections, cumulative incidence of the first event misses the full burden of disease. You might instead want the mean cumulative count, which tracks the average number of events per person over time, or a recurrent-event rate.

If the population is open and people are constantly entering and leaving, the incidence rate using person-time in the denominator handles the flux cleanly. And if you’re comparing two groups rather than describing one, the risk ratio (cumulative incidence in the exposed group divided by cumulative incidence in the unexposed group) or risk difference (one subtracted from the other) may be more informative than either group’s raw cumulative incidence alone.

Cumulative incidence remains foundational because it answers the most intuitive question a patient or policymaker can ask: “If I start in this group today, what are the chances I’ll develop this condition over the next X years?” The calculations needed to answer that question honestly range from a simple fraction to sophisticated competing-risks models, depending on how messy the real world makes your data. The key is matching the method to the messiness rather than forcing a simple formula onto complicated circumstances.

A Practical Walkthrough for Clean Data

If you’re working with a straightforward dataset and want to calculate cumulative incidence by hand or in a spreadsheet, the steps are as follows. Start by defining your cohort: who is in it, what event you’re tracking, and over what time period. Count the total number of people at risk at the start. Then count how many developed the event by the end of the observation period. Divide the second number by the first. That’s the cumulative incidence.

For example, suppose a workplace wellness program enrolls 800 employees at the start of the year and tracks who develops lower-back pain severe enough to file a workers’ compensation claim. By year’s end, 56 employees have filed such claims. Cumulative incidence is 56 divided by 800, which equals 0.07, or 7% over one year. If nobody left the company or was otherwise lost during the year, that number is exact. If 100 employees quit before the year ended, you’d need to decide whether to remove them from the denominator entirely (which slightly overstates the incidence by shrinking the denominator) or apply a life table correction that assumes they contributed partial time at risk.

That choice between the rough-and-ready approach and the more careful adjustment depends on how many people you lost and how much precision matters. For quick surveillance, dividing by the starting population is often good enough. For a peer-reviewed study or a regulatory submission, the adjustments are expected. In either case, the underlying logic is the same: you’re estimating the probability that a member of a defined group experiences a defined event within a defined period. Every complication from censoring to competing risks to dynamic populations is just a refinement of that core idea, trying to get the denominator and the numerator as honest as possible.