What Counts as Protected Health Information (PHI)?

Protected health information, commonly called PHI, is any individually identifiable health information that is created, received, maintained, or transmitted by a healthcare provider, health plan, or healthcare clearinghouse. The key phrase is “individually identifiable.” A statistic about how many people in a city have diabetes is not PHI. A record showing that you, specifically, were diagnosed with diabetes at a particular clinic on a particular date is. PHI covers a surprisingly wide range of data, from obvious items like medical records and lab results to less intuitive ones like billing addresses and device serial numbers, and its boundaries keep getting tested as health data moves into digital spaces most people never think about.

The Two Ingredients That Make Something PHI

For a piece of information to qualify as PHI under the HIPAA Privacy Rule, it needs to satisfy two conditions at once. First, it has to relate to a person’s health condition, the provision of healthcare, or payment for healthcare. Second, it has to identify the individual or provide a reasonable basis for identifying them. Strip away either ingredient and you no longer have PHI. A diagnosis code sitting in a database with no name, address, or other link to a specific person is not PHI. Your home address sitting in a phone book is not PHI. But combine a diagnosis code with a home address in a hospital’s records, and now both pieces are protected.

This two-part test is what makes PHI broader than most people expect. It is not limited to the clinical details of your visit. Payment records, insurance claim forms, appointment scheduling data, and even conversations between your doctor and a specialist about your care all count as PHI if they can be traced back to you.

The 18 HIPAA Identifiers

HIPAA’s Privacy Rule lists 18 specific types of information that can identify a person. When these identifiers are attached to health data, the combination is PHI. The list is intentionally exhaustive, designed to capture every practical way someone might be singled out:

  • Names: full name, maiden name, aliases
  • Geographic data: anything more specific than a state, including street address, city, county, and zip code (zip codes with fewer than 20,000 residents are treated as identifiers)
  • Dates: birth date, admission date, discharge date, date of death, and all ages over 89
  • Phone numbers
  • Fax numbers
  • Email addresses
  • Social Security numbers
  • Medical record numbers
  • Health plan beneficiary numbers
  • Account numbers
  • Certificate or license numbers
  • Vehicle identifiers: license plate numbers and serial numbers
  • Device identifiers and serial numbers: including implanted medical devices
  • Web URLs
  • IP addresses
  • Biometric identifiers: fingerprints, voiceprints, retinal scans
  • Full-face photographs
  • Any other unique identifying number or code

That last catch-all item is often overlooked but matters in practice. It means that any code a hospital or insurer invents to track you internally, even if it is not on the first 17 items, still counts as an identifier when paired with health data. The list is also why health data de-identification is harder than it sounds: you have to account for all 18 categories, not just the obvious ones like name and Social Security number.

How Health Data Gets De-Identified

Organizations that want to use health data for research, analytics, or quality improvement without triggering PHI protections have two official paths under HIPAA. The first is called Safe Harbor. Under Safe Harbor, you remove all 18 identifiers from the dataset and confirm that no residual information could reasonably be used to identify anyone. The second path is Expert Determination, where a qualified statistician or data scientist certifies that the risk of anyone being re-identified from the remaining data is “very small.”1PubMed Central. Beyond Safe Harbor: Automatic Discovery of Health Information De-identification Policy Alternatives

Safe Harbor is the more commonly used method because it is straightforward: follow the checklist, remove the items, and the resulting data is no longer PHI. Expert Determination is more flexible but requires hiring someone with statistical expertise, and it involves judgment calls about what “very small” risk means in a given context. In practice, Safe Harbor is the default for most hospitals and research institutions, while Expert Determination tends to appear in more complex data-sharing arrangements where Safe Harbor would strip out too much useful information.

The Re-Identification Problem

De-identification sounds like a clean solution, but the reality is messier. Even after stripping the 18 identifiers, combinations of remaining data points can sometimes be used to figure out who a record belongs to. A dataset that includes someone’s zip code, birth date, and gender, for instance, can narrow the possibilities dramatically. Research has found that re-identification vulnerability varies widely depending on geography and what other public records are available. One study estimated that the proportion of a state’s population vulnerable to unique re-identification under Safe Harbor ranged from 0.01% to 0.25%, but under the less restrictive “Limited Dataset” standard, that number jumped to between 10% and 60%.2Journal of the American Medical Informatics Association. Evaluating re-identification risks with respect to the HIPAA privacy rule

A systematic review of re-identification attacks on health data found a sharp divide between data that had been properly de-identified using standards-based methods and data that had not. When standards-based de-identification was applied, only about 0.013% of records could be correctly re-identified. When it was not, the success rate for re-identification was dramatically higher.3PLOS ONE. A Systematic Review of Re-Identification Attacks on Health Data The takeaway is that proper de-identification works well at the population level, but it is not a mathematical guarantee for every individual, especially in small or unusual populations.

Re-identification risk also depends on how many indirect identifiers remain in the dataset after the obvious ones are stripped. The more granular the remaining fields, the smaller the groups of people who share the same combination of traits, and the easier it becomes to single someone out.4PubMed Central. Evaluating re-identification risks scores in publicly available clinical trial datasets: Insights and implications This is why de-identification guidelines push organizations to think carefully about what they leave in, not just what they take out.

Genetic Data and PHI

Genetic information occupies an unusual position in the PHI landscape. Under HIPAA, your genomic data generated by a covered entity (a hospital lab running a genetic test, for example) is PHI like any other lab result. Since 2014, federal rules have granted individuals a right of access to their own laboratory test results, including genomic data.5PubMed Central. HIPAA’s Individual Right of Access to Genomic Data: Reconciling Safety and Civil Rights That means your healthcare provider generally cannot refuse to give you your genetic test results, even if the data is complex and potentially alarming without clinical context.

What makes genetic data uniquely tricky for privacy is that it is inherently identifying. Your genome is yours alone, and unlike a zip code or birth date, it cannot be generalized or rounded without losing its meaning. De-identifying a genetic dataset in any useful way is far harder than de-identifying a table of diagnoses and demographics. The Genetic Information Nondiscrimination Act of 2008, known as GINA, adds a separate layer of protection by prohibiting health insurers and employers from using genetic information to make coverage or hiring decisions, but GINA does not cover life insurance, disability insurance, or long-term care insurance. Those gaps leave real exposure for people whose genetic data is disclosed.

Direct-to-consumer genetic testing companies like 23andMe or AncestryDNA sit in a gray area. They are generally not HIPAA-covered entities because they are not healthcare providers, health plans, or clearinghouses. The privacy of your data with those companies depends on their terms of service and state-level consumer privacy laws, not on HIPAA. If you send your saliva to a consumer genetics company, your results may not qualify as PHI in the regulatory sense, even though the information is deeply personal and medically relevant.

Substance Use Disorder Records Get Extra Protection

Not all PHI is treated equally. Federal regulations under 42 CFR Part 2 impose stricter privacy protections on records related to substance use disorder treatment than HIPAA requires for other health data. These rules were originally written in the 1970s, when the stigma around addiction treatment was even more severe, and they were designed to encourage people to seek help without fear of their records being used against them in court or by employers.

In practice, 42 CFR Part 2 creates real headaches for healthcare systems trying to coordinate care. Primary care providers, hospitals, and health systems have struggled to balance good medical care with adherence to these stricter rules.6PubMed Central. Interpretation and integration of the federal substance use privacy protection rule in integrated health systems: A qualitative analysis If you are seeing both an addiction counselor and a cardiologist within the same hospital system, the substance use records may be walled off from your cardiologist unless you provide separate, specific consent for that sharing. This can result in fragmented care where one provider does not know what another is doing, potentially leading to dangerous drug interactions or duplicated treatments.

Recent regulatory changes have started to align 42 CFR Part 2 more closely with HIPAA, reducing some of the friction for integrated care. But the core principle remains: substance use treatment records carry additional protections beyond what HIPAA provides for general medical records, and those protections cannot be overridden by a standard HIPAA release form.

When Researchers Want Your Health Data

Medical research relies heavily on health data, and HIPAA does not shut the door on using PHI for research purposes. Instead, it creates a system of gatekeepers. The general rule is that a healthcare entity cannot use or disclose your PHI for research without your written authorization.7PubMed Central. Sharing clinical research data in the United States under the Health Insurance Portability and Accountability Act and the Privacy Rule But there are important exceptions.

An Institutional Review Board, or IRB, can waive the need for individual authorization if the researcher demonstrates that the privacy risk is minimal, the research cannot practically be conducted without the waiver, and PHI is genuinely necessary for the study. The researcher must also have a plan to protect the data and must destroy identifiable information as soon as it is no longer needed.8American Journal of Ophthalmology. HIPAA and Research: How Have the First Two Years Gone? This IRB waiver process is how large retrospective studies, chart reviews, and registry-based research happen without requiring consent forms from every patient whose records are included.9Ochsner Journal. Important Considerations for the Institutional Review Board When Granting Health Insurance Portability and Accountability Act Authorization Waivers

The other option for researchers is to work with fully de-identified data, which is no longer PHI at all and therefore falls outside HIPAA’s jurisdiction. Many large clinical datasets used in health services research have been stripped of identifiers for exactly this reason. The tradeoff is that de-identification can reduce the data’s usefulness, particularly for studies that need to track patients over time or link records across different care settings.

Public Health Reporting and Surveillance

HIPAA includes explicit exceptions that allow PHI to be shared for public health purposes without patient consent. These exceptions are what make disease surveillance, outbreak investigation, and mandatory reporting of conditions like tuberculosis or HIV possible. Without them, public health agencies would be unable to monitor population health in any meaningful way.

The Privacy Rule provides several mechanisms for this, including the use of existing public health laws and regulations, allowances for public health activities, de-identification, research waivers, and limited datasets with aggregate reporting.10Journal of the American Medical Informatics Association. A Model for Expanded Public Health Reporting in the Context of HIPAA In practical terms, your doctor can report your positive tuberculosis test to the local health department without asking you first, and the hospital can share data with the CDC during a pandemic, because these activities are carved out as permissible uses of PHI under the law.

These exceptions are narrowly defined. A public health authority can receive PHI for surveillance and investigation, but it cannot turn around and use that data for marketing or employment decisions. The data stays within the public health function it was shared for.

Tracking Pixels and the Digital Frontier

One of the most consequential developments in PHI regulation has nothing to do with paper medical records. It involves the tiny, invisible tracking pixels embedded on hospital and health plan websites. These snippets of code, typically placed by third-party companies for advertising analytics, can transmit data about what pages a visitor views, what they search for, and what forms they fill out. When a hospital website uses third-party tracking pixels, the data flowing to outside companies can include information that qualifies as PHI, such as which condition-specific pages a logged-in patient visited or what appointment types they scheduled.

Research has found that hospitals using third-party tracking pixels experience a meaningfully higher probability of data breaches. One study estimated that pixel use increases breach probability by at least 1.4 percentage points, which may sound small until you consider that the average breach probability in the sample was about 3%. That represents roughly a 46% relative increase in breach risk. The same study found that hospitals with third-party pixels saw a 13% increase in unintended data disclosures. First-party pixels, which keep data within the hospital’s own systems, showed no such relationship, suggesting that the vulnerability is specifically about sending data to external parties, not about using tracking technology in general.11PubMed Central. Beyond the click: Pixel tracking technologies and patient data security in hospitals

This became a major enforcement issue starting in 2022, when the U.S. Department of Health and Human Services warned that regulated entities using tracking technologies on their websites or apps could be violating HIPAA if those tools transmitted PHI to third parties without proper authorization. Several major health systems subsequently disclosed that they had been sharing patient data through tracking pixels for years, often without realizing it. The lesson for patients is that PHI exposure can happen through channels that have nothing to do with a doctor’s office or a misplaced paper chart.

Billing Data and Insurance Claims

Health insurance claims are PHI. This catches people off guard because billing feels administrative rather than medical, but a single insurance claim can reveal a diagnosis code, a procedure performed, the treating provider’s identity, the dates of service, and the patient’s name and insurance ID. All of that is individually identifiable health information.

Linking insurance claims data with electronic medical records is common in health services research and quality improvement, but it creates its own complications. One study that attempted to link insurer claims with medical records using a business associate agreement found that when the linked dataset was restricted to patients with at least one medication and one diagnosis in the evaluation year, about 90% of the linked population was lost.12PubMed Central. The challenges of linking health insurer claims with electronic medical records The privacy protections worked, in the sense that they prevented easy, broad matching, but the result was a dataset so narrow it was hard to use for its intended purpose. This tension between data utility and privacy protection plays out constantly in healthcare analytics.

For you as a patient, the practical implication is that your Explanation of Benefits statements, claim denials, and even the fact that a claim was submitted at all are PHI. If your employer self-funds its health plan, the company itself is subject to HIPAA restrictions on how it handles claims data, even though it is also your employer. Firewalls are supposed to exist between the benefits administration side and the human resources side, but the potential for conflicts of interest is built into the structure.

Common Misconceptions About What PHI Covers

Several widespread misunderstandings about PHI trip people up regularly. The first is that HIPAA applies to everyone who handles health information. It does not. HIPAA’s Privacy Rule applies to “covered entities,” which are healthcare providers that transmit information electronically, health plans, and healthcare clearinghouses, plus their “business associates,” which are companies that handle PHI on their behalf. Your fitness app, your employer (unless it runs a self-funded health plan), and your neighbor who overheard your diagnosis in a waiting room are not covered by HIPAA. They may be bound by other privacy laws, but not HIPAA.

The second misconception is that all health-related information is PHI. It is not. Health data only becomes PHI when it is held by or created by a covered entity and is individually identifiable. The blood pressure reading from your home cuff that you jot in a personal notebook is not PHI. The same reading entered into your medical record by a nurse at a clinic is.

A third misunderstanding is that HIPAA prevents your doctor from talking to your family about your care. In many situations, healthcare providers can share information with family members involved in your care unless you specifically object. HIPAA allows for reasonable professional judgment in these situations, which is why a surgeon can update your spouse after an operation without violating the law, as long as you have not asked them not to.

PHI After Death

HIPAA protections do not evaporate when someone dies. PHI of deceased individuals remains protected for 50 years after the date of death. During that period, the same rules about use and disclosure apply, with some additional provisions. Personal representatives of a deceased person’s estate generally have the same rights to access the decedent’s PHI as the person would have had while alive. This means an executor can request medical records from a hospital, and the hospital is required to provide them under the same access rules that apply to living patients.

The 50-year window is long enough to cover most practical scenarios, but it also means that historical medical records from the mid-20th century onward may still carry PHI protections. Researchers working with archival health data sometimes run into this limitation, particularly when studying conditions where the patient population was small enough that even basic demographic details could identify someone.

Psychotherapy Notes Stand Apart

Within the universe of PHI, psychotherapy notes occupy a uniquely privileged position. These are the personal notes a mental health professional takes during or after a counseling session, kept separate from the rest of the medical record. HIPAA gives psychotherapy notes stronger protection than other PHI. They cannot be disclosed for most purposes, including treatment, payment, and healthcare operations, without the patient’s specific authorization. Even the IRB waiver process that allows researchers to access other types of PHI without consent does not automatically extend to psychotherapy notes.

This distinction matters because psychotherapy notes often contain the most sensitive information a person shares in a healthcare setting. The extra layer of protection exists specifically because people might not speak candidly in therapy if they believed their therapist’s notes could be pulled into an insurance review or shared with other providers without their knowledge. Progress notes and treatment summaries that go into the regular medical chart, by contrast, are treated like any other PHI and can be shared under the standard HIPAA rules for treatment and payment.