How to Test for Depression: Screening Tools & Symptoms

Depression screening relies on short questionnaires that ask about mood, sleep, energy, and appetite over the past two weeks, and the most common of these, the Patient Health Questionnaire-9 (PHQ-9), can be completed in under five minutes. But a screening tool is not a diagnosis. Every widely used depression questionnaire is designed to flag people who should be evaluated further, not to deliver a verdict on its own. Understanding what these tools actually measure, where they work well, and where they fall short matters if you are trying to figure out whether what you are feeling qualifies as clinical depression.

The PHQ-9 and How It Works

The PHQ-9 is the screening instrument you are most likely to encounter in a doctor’s office, an emergency department, or even an online mental health portal. It consists of nine questions, each mapping to one of the core symptoms used to diagnose major depressive disorder. You rate each item from 0 (“not at all”) to 3 (“nearly every day”), producing a total score between 0 and 27. A score of 10 or above is the most commonly used threshold for a positive screen.

That cutoff of 10 has held up well across large bodies of evidence. A major meta-analysis pooling individual participant data from 58 studies and over 17,000 people found that a score of 10 or higher produced a sensitivity of 0.88 and a specificity of 0.85 when measured against a semistructured clinical interview.1BMJ. Accuracy of Patient Health Questionnaire-9 (PHQ-9) for screening to detect major depression: individual participant data meta-analysis In plain terms, the test correctly identified roughly 88% of people who had major depression and correctly cleared about 85% of people who did not. The original validation study of the PHQ-9 found nearly identical numbers: 88% sensitivity and 88% specificity at the same cutoff.2PubMed Central. The PHQ-9: validity of a brief depression severity measure

Those numbers are good for a screening tool, but they are not perfect. A separate systematic review of 42 studies in primary care found that PHQ-9 sensitivity ranged from as low as 0.37 to as high as 0.98 depending on the study population and how depression was confirmed.3PubMed. Screening for depression in primary care with Patient Health Questionnaire-9 (PHQ-9): A systematic review In some settings, positive predictive values were as low as 9%, meaning that in populations with low depression rates, most people who scored above 10 did not actually have major depression when formally evaluated. This is one of the most misunderstood aspects of screening: a “positive” result on the PHQ-9 means you warrant a proper clinical assessment, not that you have been diagnosed.

What Happens After a Positive Screen

The US Preventive Services Task Force recommends universal depression screening for all adults, including pregnant and postpartum individuals and older adults, with the important caveat that screening should only happen where systems are in place for accurate diagnosis, effective treatment, and follow-up.4PubMed. Screening for Depression and Suicide Risk in Adults: US Preventive Services Task Force Recommendation Statement That caveat is doing a lot of work. Screening without a plan for what comes next can lead to people being told they are depressed, prescribed medication, and never seen by a mental health professional.

In ideal practice, a positive screen triggers a second stage: a longer, structured or semistructured clinical interview conducted by a trained provider. This two-stage approach is how the PHQ-9 was designed to be used. The clinician explores the duration, severity, and context of your symptoms, rules out medical causes like thyroid dysfunction or medication side effects, and checks for other conditions that can look like depression. Only after that evaluation does a formal diagnosis get made using standard criteria.

The gap between ideal and actual practice is wide. In busy primary care clinics, the PHQ-9 score alone sometimes becomes a de facto diagnosis, particularly at higher scores. If you score above 10 and your doctor immediately suggests medication without asking many follow-up questions, you are within your rights to ask for a more thorough evaluation or a referral.

Severity Tiers on the PHQ-9

Beyond the simple positive-or-negative cutoff, the PHQ-9 breaks scores into severity categories: 5 to 9 is mild, 10 to 14 is moderate, 15 to 19 is moderately severe, and 20 to 27 is severe. These tiers loosely correspond to different levels of clinical urgency. A meta-analysis examining various cutoff thresholds found that sensitivity and specificity remained fairly stable across cutoff scores from 8 to 11, but specificity jumped substantially at higher cutoffs, reaching about 0.96 at a cutoff of 15.5PubMed Central. Optimal cut-off score for diagnosing depression with the Patient Health Questionnaire (PHQ-9): a meta-analysis In other words, the higher your score, the more confident a clinician can be that something clinically significant is happening.

Some stepped-care models use these tiers to guide treatment. A score in the mild range might trigger watchful waiting or a recommendation for self-help resources. Moderate scores might prompt a referral for psychotherapy. Scores in the severe range often lead to a combination of medication and therapy, with closer follow-up. The tiers are rough guides rather than rigid decision rules, but they help clinicians triage when resources are limited.

Self-Report Tools Versus Clinician-Rated Scales

The PHQ-9 is a self-report measure: you fill it out yourself. That is convenient, but it introduces a specific kind of bias. People who are more introverted or who tend toward self-criticism often rate their symptoms more severely on self-report tools than a clinician would during an interview. Research comparing the Beck Depression Inventory (a self-report tool) with the Hamilton Depression Rating Scale (a clinician-rated tool) found that personality traits like high neuroticism and low extraversion were associated with higher self-reported symptom scores relative to clinician ratings.6PubMed. Sensitivity to detect change and the correlation of clinical factors with the Hamilton Depression Rating Scale and the Beck Depression Inventory in depressed inpatients The two types of tools are best thought of as complementary. One captures how you experience your own symptoms; the other captures what a trained observer sees.

The Hamilton scale also tends to show larger reductions in scores over the course of treatment, which is one reason it is the preferred outcome measure in clinical trials of antidepressants.6PubMed. Sensitivity to detect change and the correlation of clinical factors with the Hamilton Depression Rating Scale and the Beck Depression Inventory in depressed inpatients If you are tracking your own recovery with a self-report questionnaire, your scores might seem to improve more slowly than what your clinician observes. That discrepancy does not mean one of you is wrong; it reflects the different lenses the tools use.

Screening During Pregnancy and Postpartum

The Edinburgh Postnatal Depression Scale (EPDS) is the most widely used screening tool for depression during and after pregnancy. It was originally developed for postpartum use but is now routinely given to pregnant individuals as well. A large meta-analysis of individual participant data from 58 studies found that a cutoff of 11 or higher produced a sensitivity of 0.81 and a specificity of 0.88 when validated against semistructured interviews.7PubMed Central. Accuracy of the Edinburgh Postnatal Depression Scale (EPDS) for screening to detect major depression among pregnant and postpartum women: systematic review and meta-analysis of individual participant data A slightly lower cutoff of 10 boosted sensitivity to 0.85 while keeping specificity at 0.84, and accuracy was similar for pregnant and postpartum women.

One distinctive feature of the EPDS is item 10, which asks about self-harm thoughts. A version of the scale with that item removed (the EPDS-9) correlates nearly perfectly with the full version and performs similarly in identifying depression.8PubMed. The screening accuracy of the Edinburgh Postnatal Depression Scale (EPDS) to detect perinatal depression with and without the self-harm item in pregnant and postpartum women However, neither version was particularly good at identifying people who ended up needing medication; sensitivity for that outcome peaked at only about 58% even at the most favorable cutoff. The EPDS catches depression that meets screening criteria, but it is not designed to predict treatment needs.

A practical concern for non-English-speaking populations is that translated versions of the EPDS often perform worse. A systematic review of 14 locally validated versions used in low- and lower-middle-income countries found that none met the recommended validation standard of at least 80% sensitivity, specificity, and positive predictive value simultaneously.9PubMed Central. Reliability and validity of the Edinburgh Postnatal Depression Scale (EPDS) for detecting perinatal common mental disorders (PCMDs) among women in low-and lower-middle-income countries: a systematic review Compromises made during translation and cultural adaptation likely explain much of this drop.

Screening in Older Adults

Depression in older adults frequently gets missed because symptoms overlap with normal aging, grief, medical illness, and medication side effects. The Geriatric Depression Scale (GDS) was designed specifically for this population, using a yes-or-no format rather than graded responses, which makes it easier for older adults to complete. A meta-analysis comparing the GDS with the Cornell Scale for Depression in Dementia (CSDD) found that in older adults without dementia, the GDS performed well, with pooled sensitivity of 0.88 and specificity of 0.82.10PubMed. Which of the Cornell Scale for Depression in Dementia or the Geriatric Depression Scale is more useful to screen for depression in older adults?

The critical limitation is cognitive impairment. In people with Alzheimer’s disease, the GDS essentially becomes unreliable. One study found that the GDS produced an area under the curve of just 0.66 in an Alzheimer’s group, which is barely better than flipping a coin, compared with 0.85 in cognitively intact older adults.11PubMed. Use of the Geriatric Depression Scale in dementia of the Alzheimer type For people with dementia, the CSDD, which relies on caregiver input alongside direct observation, is the better instrument, with pooled sensitivity of 0.91 in the dementia subgroup.10PubMed. Which of the Cornell Scale for Depression in Dementia or the Geriatric Depression Scale is more useful to screen for depression in older adults? If you are trying to assess depression in an older family member with memory problems, a self-report tool is not the right approach.

Adolescent Screening

For adolescents, a modified version of the PHQ-9 (the PHQ-9A or PHQ-9M) is increasingly used during annual wellness visits. The modification adds language relevant to younger populations while keeping the same nine-item structure. Implementation studies have shown that incorporating universal screening during well visits for 12- to 18-year-olds is feasible and effective at identifying depression that would otherwise go undetected.12PubMed. Implementation of Universal Adolescent Depression Screening: Quality Improvement Outcomes The consensus from primary care implementation research is that the PHQ-9M works, but only when staff and providers are trained to act on positive screens.13PubMed. Universal Depression Screening for Pediatric Adolescent Patients at Well Visits Using the PHQ-9M Within the Pediatric Primary Care Clinic

A common stumbling block in adolescent screening is that teenagers may not describe their experience the way the questionnaire expects. Irritability, anger, and somatic complaints like headaches or stomach pain are often more prominent than the classic “sad mood” that screening questions emphasize. If a teen scores below the threshold but their behavior and functioning tell a different story, the screen should not be taken as the final word.

Why Standard Questionnaires Miss Some People

Most widely used depression questionnaires were developed in Western, English-speaking settings, and research has exposed meaningful cross-cultural differences in how depression manifests on these instruments. A study comparing Beck Depression Inventory scores across six countries found that people in Finland endorsed more irritability and sleep disruption, people in Mexico endorsed more self-criticalness and feelings of punishment, and people in Japan endorsed more sadness, all using the same questionnaire.14PubMed Central. Cross-cultural comparison of depressive symptoms on the Beck Depression Inventory-II, across six population samples These are not trivial differences. A questionnaire that weights certain symptom clusters heavily will be more sensitive in populations where those clusters dominate and less sensitive where they do not.

An alternative approach has been to build instruments from the ground up using symptom presentations from multiple cultures. Researchers developing the International Depression Symptom Scale found that it predicted functional impairment slightly but significantly better than the PHQ-9 in at least one non-Western sample, suggesting that the standard Western model of depression misses locally relevant expressions of distress that are relevant across cultures, not just in one location.15PubMed Central. Development and cross-cultural testing of the International Depression Symptom Scale (IDSS): a measurement instrument designed to represent global presentations of depression This work is still in its early stages, but it underlines a real limitation: if you come from a cultural background where sadness is expressed through bodily symptoms, social withdrawal, or idioms that do not translate neatly into questionnaire items, a standard screening tool may underestimate what you are going through.

The Suicide Risk Question and Its Limits

Item 9 on the PHQ-9 asks how often you have been bothered by “thoughts that you would be better off dead, or of hurting yourself.” It is the single most anxiety-producing question on the form for both patients and providers. And it turns out to be a poor standalone measure of actual suicide risk.

In a validation study comparing item 9 against the Columbia Suicide Severity Rating Scale, about 41% of patients endorsed item 9, but only about 13% were identified as genuinely at risk by the more comprehensive scale. Sensitivity for suicide risk was reasonable at about 88%, but the positive predictive value was just 29%, meaning most positive responses on item 9 did not correspond to clinically meaningful suicide risk.16PubMed. The PHQ-9 Item 9 based screening for suicide risk: a validation study of the Patient Health Questionnaire (PHQ)-9 Item 9 with the Columbia Suicide Severity Rating Scale (C-SSRS) A study in coronary artery disease patients was even more striking: of 110 patients who endorsed item 9, only about 20% reported actual suicidal thoughts, and only about 8% described a suicide plan in the past year.17Journal of Psychosomatic Research. The PHQ-9 versus the PHQ-8 – Is item 9 useful for assessing suicide risk in coronary artery disease patients? Data from the Heart and Soul Study

The flip side is also concerning. Among medical inpatients who screened positive for suicide risk on a dedicated scale, about 63% had not endorsed item 9 at all, and roughly 31% had screened entirely negative for depression on the PHQ-9.18PubMed Central. Limitations of Screening for Depression as a Proxy for Suicide Risk in Adult Medical Inpatients Suicide risk and depression overlap, but they are not the same thing, and a depression screening tool is not a substitute for a proper suicide risk assessment. If you or someone you know is having thoughts of suicide, a full evaluation with a trained clinician is warranted regardless of what any questionnaire says.

Ruling Out Bipolar Disorder

One trap in depression screening is that the depressive phase of bipolar disorder looks identical to unipolar depression on any standard questionnaire. The PHQ-9 cannot tell the difference, and this matters because the treatment is different: antidepressants without a mood stabilizer can trigger mania in someone with bipolar disorder. The Mood Disorder Questionnaire (MDQ) is a brief self-report tool designed to flag bipolar disorder, asking about lifetime experiences of manic or hypomanic symptoms.19PubMed Central. The Mood Disorder Questionnaire: A Simple, Patient-Rated Screening Instrument for Bipolar Disorder

However, the MDQ has its own limitations. In patients who presented with depression and had no prior bipolar diagnosis, the standard MDQ scoring yielded a sensitivity of just 0.29, meaning it missed about 70% of bipolar cases.20PubMed. Bipolarity in depressive patients without histories of diagnosis of bipolar disorder and the use of the Mood Disorder Questionnaire for detecting bipolarity A modified scoring approach (dropping the requirement that symptoms co-occur and cause functional impairment) improved sensitivity to 0.68 at the cost of lower specificity. The takeaway is that no brief questionnaire reliably separates unipolar from bipolar depression. If a clinician suspects bipolarity based on your history, a thorough clinical interview is the only reliable path.

Screening When You Have a Chronic Illness

Chronic illness and depression frequently coexist, but the overlap in symptoms creates a measurement headache. Fatigue, poor appetite, sleep problems, and trouble concentrating are symptoms of both depression and conditions like diabetes, heart disease, and HIV. This means the PHQ-9 can give inflated scores in people with chronic medical conditions because some of the items are picking up physical illness, not mood disorder.

A study of chronic care patients in primary care found that the PHQ-9 maintained reasonable overall accuracy (area under the curve of 0.85) in patients with HIV and hypertension, but at the commonly used cutoff of 9, sensitivity dropped to just 51%.21PubMed Central. The validity of the Patient Health Questionnaire for screening depression in chronic care patients in primary health care in South Africa In other words, it missed about half of the people who actually had major depression. Some researchers have suggested that the WHO-5 Well-Being Index, which focuses on positive mood and vitality rather than symptom counts, may be a better initial screen for people with chronic illness because its items are less likely to overlap with physical symptoms.22PubMed Central. Rapid Screening of Psychological Well-Being of Patients with Chronic Illness: Reliability and Validity Test on WHO-5 and PHQ-9 Scales

Melancholic, Atypical, and the Problem of Subtypes

Depression is not one thing. Clinicians sometimes distinguish between melancholic depression (marked by a pervasive inability to feel pleasure, worsening mood in the morning, psychomotor changes, and appetite loss) and atypical depression (characterized by mood that lifts in response to good events, increased sleep, weight gain, and a heavy feeling in the limbs). These subtypes are recognized in diagnostic manuals, and there have been efforts to build screening tools that can differentiate them.23PLoS ONE. Clinical Patterns and Treatment Outcome in Patients with Melancholic, Atypical and Non-Melancholic Depressions

A clinician-rated measure called the Sydney Melancholia Prototypic Index performed well in separating melancholic from non-melancholic depression, with sensitivity of 0.84 and specificity of 0.92 in one study.24PubMed. Discriminating melancholic and non-melancholic depression by prototypic clinical features But a deeper analysis of symptom profiles in a large dataset found that the melancholic and atypical labels did not actually reduce the enormous variability in how depression looks from person to person. Most symptom profiles, even within supposed subtypes, contained five or fewer individuals, meaning almost everyone’s depression was unique.25PubMed Central. Heterogeneity in major depression and its melancholic and atypical specifiers: a secondary analysis of STAR*D The evidence is honest in a humbling way: depression categories exist on paper, but the lived reality is that your particular mix of symptoms is unlikely to match anyone else’s exactly.

Blood Tests and Biomarkers for Depression

There is no blood test for depression, despite decades of research. Hundreds of candidate biomarkers have been studied, spanning inflammation markers, stress hormones, neurotransmitter metabolites, and brain-derived growth factors. None has proven reliable enough to use in clinical practice for diagnosis or treatment selection.26PubMed Central. Biomarkers for depression: recent insights, current challenges and future prospects The obstacle is not that biology is irrelevant but that depression is biologically heterogeneous. Two people meeting the same diagnostic criteria can have completely different underlying biology, which means a single biomarker, or even a panel, cannot capture the full picture.27Psychiatry Investigation. Biomarkers of Major Depressive Disorder: Knowing is Half the Battle

Newer lines of investigation are exploring digital biomarkers instead: changes in voice pitch and speech rhythm detectable through a smartphone microphone. A scoping review of this research found that prosodic voice features show promise for monitoring depressive symptoms, though their standalone predictive power has not been confirmed, and combining voice data with other sensor inputs like sleep patterns and physical activity improves accuracy.28PubMed Central. Use of voice features from smartphones for monitoring depressive disorders: Scoping review These tools are still research-grade, not clinical-grade, but they represent a genuinely different approach to the measurement problem: passive, continuous monitoring rather than a snapshot questionnaire administered once at a clinic visit.