The FDA’s Sentinel System is a large-scale electronic surveillance network that monitors the safety of drugs, vaccines, and other medical products after they reach the market. Rather than waiting for doctors and patients to voluntarily report problems, Sentinel actively queries health data from insurance claims, hospital records, and pharmacy dispensings across multiple institutions, covering roughly 700 million person-years of observation time nationwide. It was built in response to a 2007 law that required the FDA to move beyond its traditional reliance on passive reporting, and it has since become one of the largest post-market safety systems in the world.
Why the FDA Needed Something New
For decades, the primary way the FDA learned about drug side effects after approval was through spontaneous reporting. A patient experienced an unexpected reaction, a physician filed a report, and the FDA’s MedWatch system collected it. This approach has an obvious flaw: most adverse events go unreported. Estimates vary, but only a small fraction of real-world drug problems ever make it into the voluntary reporting database. That means signals can take years to surface, and when they do, there is no reliable denominator to calculate how common a problem actually is.
Regulators around the world recognized this limitation and began exploring more proactive approaches that use existing health data rather than waiting for individual reports to trickle in. The FDA took a major step with the FDA Amendments Act of 2007, which required the agency to establish a system for monitoring risks associated with drug and biologic products using data from multiple sources. That legislative mandate gave birth to what eventually became the Sentinel System.
From Mini-Sentinel to the Full System
The FDA did not build Sentinel all at once. It started with a pilot program called Mini-Sentinel, launched in 2009, to test whether it was even feasible to build a large, multi-organizational data network where each participating institution kept control of its own data. The pilot demonstrated that yes, this kind of distributed system could work for active safety surveillance. By 2016, the FDA transitioned from the pilot into the fully operational Sentinel System, expanding both the number of data partners and the analytical capabilities available.
Today, Sentinel draws on electronic health data from a network of data partners that includes health insurers, health systems, and academic medical centers. The data covers administrative claims, outpatient pharmacy dispensings, laboratory results, and increasingly, information from electronic health records. The system’s reach, covering about 700 million person-years of longitudinal health data, gives it the statistical power to detect safety signals that would be invisible in smaller datasets.
The Distributed Data Model
The architectural choice at the heart of Sentinel is the distributed data model, and it solves two problems at once: privacy and speed. In a centralized system, all participating hospitals and insurers would need to send their patient-level data to a single FDA-controlled warehouse. That would raise enormous privacy concerns and create a massive target for data breaches. Sentinel avoids this by leaving the data where it already lives. Each data partner keeps its own records behind its own firewall. When the FDA wants to investigate a potential safety signal, it sends a standardized computer query out to each partner. The partners run the query locally, against their own data, and send back only aggregated, summary-level results. No individual patient records leave the partner’s system.
This design was validated during the Mini-Sentinel pilot, which showed that a large, multi-organizational distributed network where organizations retain possession of their data could function as an active surveillance system. The approach has since become a model for health data networks beyond drug safety.
How Data Gets Standardized
Different hospitals code diagnoses differently. Insurance claims look nothing like electronic health records. Pharmacy dispensing data has its own format. For Sentinel to run a single query across all its partners and get comparable results, everyone’s data needs to speak the same language. That is the job of the Sentinel Common Data Model.
Each data partner transforms its raw source data into this standardized format. The Common Data Model uses a person-centered design, organizing information around individual patients with attributes for demographics, drug exposures, and medical conditions. Drug “eras” are constructed to represent periods when a patient was consistently using a medication, drawn from pharmacy dispensings, prescriptions, and other medication history. Condition “eras” aggregate diagnoses that occurred within a single episode of care. Both drugs and conditions are mapped to standardized medical vocabularies so that, for instance, every partner’s version of “heart failure” points to the same concept regardless of which coding system the original data used. Once the transformation is complete, the data undergoes rigorous quality checks before it can be used for any Sentinel analysis.
The Querying Tools
Sentinel does not require a team of programmers to write custom code every time the FDA wants to investigate a potential drug risk. Instead, it relies on a suite of pre-tested, standardized computer programs that can be configured for different study questions. These tools allow analysts to perform sophisticated analyses, from simple descriptive counts to complex inferential studies, without exchanging individual-level data across sites.
The tools are modular and publicly available. In fact, researchers in Taiwan demonstrated that Sentinel’s querying tools could be applied directly to their own national health insurance database after it had been formatted into the Sentinel Common Data Model, producing comparable results to U.S. studies. That kind of portability speaks to how well-designed the underlying software is: the same analytical programs work across entirely different health care systems because the data format is consistent.
What Kinds of Studies Sentinel Can Run
Sentinel supports several types of safety assessments, and the level of sophistication has grown considerably since the early pilot days.
At the simpler end, the system can run descriptive queries that characterize how a drug is being used in the real world: how many people are taking it, for how long, at what doses, and what conditions they have. These are useful for understanding whether a drug is being prescribed as intended or drifting into off-label territory.
At the more complex end, Sentinel can conduct comparative cohort studies that approximate the kind of evidence you would get from a clinical trial, but using real-world data from millions of patients. A good example is a prospective surveillance study that compared outcomes between new users of the blood thinner rivaroxaban and new users of warfarin in patients with atrial fibrillation. That study sequentially matched over 36,000 rivaroxaban users to nearly 80,000 warfarin users using propensity score methods and tracked outcomes over multiple monitoring periods. The sequential design meant the FDA could detect emerging safety signals in near-real time rather than waiting for a years-long study to finish.
Another demonstration involved the diabetes drug glyburide, where researchers used Sentinel’s propensity score matching tool to see if it could reproduce a well-known association between glyburide and serious hypoglycemia compared to a similar drug, glipizide. Using different propensity score approaches, the tool consistently detected the expected increased risk for glyburide, with hazard ratios ranging from about 1.36 to 1.51 depending on the matching method. The point was not to discover something new but to confirm that the system’s automated tools could reliably and quickly replicate known safety findings.
Privacy Protections and Governance
Because Sentinel touches health data from millions of Americans, the system has a layered governance structure designed to keep individual information protected. A dedicated privacy panel advises on the protection of personal health information, and the system’s operating principles are built around fair information practices: protecting privacy, maintaining data security and integrity, guarding the confidentiality of proprietary information from data partners, and preventing conflicts of interest. The distributed architecture itself is the strongest privacy safeguard, since patient-level data never leaves the institution that holds it. The FDA only ever receives aggregate, de-identified results.
Sentinel During COVID-19
The pandemic put Sentinel’s flexibility to the test. Early in 2020, the system had to pivot quickly to support a range of FDA data needs that nobody had anticipated when the system was designed. The response involved multiple tracks: incorporating new data sources, creating a rapidly refreshed database that could keep pace with a fast-moving outbreak, developing protocols to study the natural history of COVID-19 itself, and validating algorithms for identifying COVID-19 patients in administrative claims data. Sentinel also coordinated with other national and international surveillance initiatives to ensure its work complemented rather than duplicated what other groups were doing.
This rapid adaptation mattered because the FDA was simultaneously overseeing Emergency Use Authorizations for vaccines, treatments, and diagnostics. Having a national electronic health data network that could be retooled within weeks, rather than the months or years a new study would take to stand up, gave regulators a tool they would not have had a decade earlier.
Going Beyond Claims Data With Electronic Health Records
Sentinel was originally built primarily on administrative claims data, the billing records generated when you visit a doctor or fill a prescription. Claims data is great for scale and completeness of certain fields like medication dispensings, but it misses a lot. It does not capture what the doctor wrote in your chart notes, what your lab values were, or whether a diagnosis was confirmed or just considered and ruled out.
Increasingly, Sentinel has been integrating electronic health record data to fill those gaps. One area where this has paid off is vaccine surveillance. A study within the FDA’s Biologics Effectiveness and Safety Initiative used natural language processing to search clinical notes for vaccine administrations that were not captured in structured data fields. Of the total vaccine administrations identified, about 86 percent appeared in structured data, but the NLP algorithm found an additional 459 administrations that structured data alone missed, representing a roughly 17 percent increase in detection. When those NLP-identified administrations were validated, over 96 percent were confirmed as definite vaccine administrations, with more than a third found solely in the unstructured notes and nowhere in the billing data.
That kind of improvement matters for vaccine safety monitoring, where completeness of exposure data directly affects the system’s ability to detect rare adverse events. If you do not know who received a vaccine, you cannot assess whether something bad happened at a higher-than-expected rate after vaccination.
Artificial Intelligence and the Future of Sentinel
The FDA has signaled that the next phase of Sentinel’s evolution will involve deeper integration of AI and machine learning. The Sentinel Innovation Center has been working on incorporating generative AI and machine learning into the surveillance workflow, particularly for tasks that currently require manual expert effort: identifying relevant patient populations in unstructured text, improving the accuracy of outcome definitions, and potentially flagging emerging safety signals faster than traditional epidemiologic methods can.
The approach is not to replace human judgment but to combine human and artificial intelligence for a more robust system. Post-market safety surveillance involves judgment calls at every stage, from how you define a medical outcome to how you handle confounding variables to how you interpret borderline results. AI can help with pattern recognition and data extraction at a scale no team of epidemiologists could manage manually, but the decisions about what those patterns mean still require people who understand both the medicine and the regulatory context.
Real-World Regulatory Impact
Sentinel is not just a research exercise; its findings feed directly into FDA regulatory decisions. One concrete example involves proton pump inhibitors, the widely used class of heartburn and acid reflux drugs. In 2010, the FDA required a class-wide label change warning about the risk of fractures with prolonged PPI use. Sentinel’s tools were used to assess whether that label change actually affected prescribing behavior and patient outcomes. The analysis covered nearly 1.5 million incident PPI users in the period before the label change and over 2.2 million in the period after. Users with a year or more of exposure decreased modestly from about 8.4 to 7.5 percent, and average days of PPI supplied per user dropped as well. The proportion of patients with fractures also decreased, from 4.4 to 3.1 percent, though osteoporosis screening and interventions did not appear to increase.
Studies like this are valuable not just for confirming that a regulatory action had some effect, but for understanding how the effect played out in practice. A label change is only useful if doctors and patients actually respond to it, and Sentinel provides the data infrastructure to measure that response across millions of people rather than relying on surveys or small studies.
How Sentinel Differs From Clinical Trials
Clinical trials remain the gold standard for establishing whether a drug works. But they have well-known limitations for safety monitoring. Trials typically enroll a few thousand carefully selected patients, follow them for a defined period, and exclude people with complicated medical histories or multiple medications. The real world is messier. Drugs get used by older, sicker, more diverse populations, for longer durations, and in combinations that no trial tested. Rare side effects that affect one in ten thousand users may never appear in a trial of five thousand people.
Sentinel fills this gap by monitoring what happens after a drug enters routine clinical use. It cannot prove causation with the same rigor as a randomized trial, because it relies on observational data where patients were not randomly assigned to one treatment or another. But the system’s analytical tools, particularly the propensity score methods that attempt to account for differences between patient groups, bring the evidence closer to causal inference than raw observational data would allow. The tradeoff is a deliberate one: less internal validity than a trial, but far greater scale, diversity, and real-world relevance.
International Influence
Sentinel’s design has influenced how other countries approach post-market surveillance. The distributed data model, where institutions keep their data and respond to standardized queries, has been adopted or adapted by regulatory systems in other countries. As noted, researchers in Taiwan formatted their national health insurance database to match the Sentinel Common Data Model and successfully applied Sentinel’s querying tools to answer the same study questions that had been investigated in the U.S. The ability to replicate analyses across entirely different health care systems using the same tools opens the door to international safety surveillance collaborations that would have been impractical with bespoke, country-specific analytical systems.
Other countries have pursued their own versions of active surveillance, but Sentinel’s combination of scale, standardized tools, and a governance model that has navigated the complex privacy landscape of the U.S. health care system makes it a reference point for regulators worldwide. The underlying philosophy, that existing health data can be harnessed for ongoing safety monitoring without requiring a massive centralized database, has proven broadly applicable regardless of how a country’s health care system is organized.
What Sentinel Does Not Do
For all its capabilities, Sentinel has boundaries worth understanding. It does not monitor every medical product on the market in real time. The FDA uses it to investigate specific safety questions, either proactively (routine surveillance of newly approved drugs) or reactively (following up on signals from spontaneous reports or published literature). It is a tool the FDA deploys for particular investigations, not an automated alarm system that rings whenever something goes wrong.
Sentinel also cannot detect problems that do not leave a trace in billing records or electronic health charts. If a side effect is subtle, does not prompt a medical visit, or gets coded under a vague diagnosis, it may not show up in the data. The system is strongest for outcomes that reliably generate medical encounters and diagnostic codes: hospitalizations, emergency department visits, filled prescriptions, documented lab results. It is weaker for patient-reported outcomes like fatigue or mood changes that may never make it into a medical record.
Finally, Sentinel’s data has a built-in lag. Claims data can take weeks to months to be processed and made available. While the system created a rapidly refreshed database during COVID-19, under normal operations there is a delay between a patient’s medical encounter and when that encounter appears in the data available for querying. For fast-moving safety crises, the FDA still relies on other mechanisms, including its traditional adverse event reporting system, to get early warnings before Sentinel data catches up.