The Mood Disorder Questionnaire uses a three-part scoring system: you check yes or no on 13 symptom items, indicate whether several of those symptoms happened during the same time period, and rate how much functional trouble those symptoms caused. A screen counts as positive when all three conditions are met: seven or more “yes” answers on the symptom items, a “yes” to the co-occurrence question, and at least a “moderate” or “serious” rating on the impairment question. That standard algorithm is straightforward on paper, but the details of how each part interacts, where the cutoffs come from, and when the standard rules should be adjusted are worth understanding.
The Three Parts of the Questionnaire
The MDQ was designed as a brief, self-administered screen that takes roughly five minutes to complete. It is not a diagnostic tool on its own; it flags people who should be evaluated further for bipolar disorder. The questionnaire has three distinct sections, and all three feed into the final score.
The first section is a list of 13 yes/no items describing experiences associated with mania or hypomania. These cover a wide range: feeling unusually energetic, needing much less sleep, talking faster than normal, having racing thoughts, being much more easily distracted, taking on many more projects than usual, being more social or outgoing, being more interested in sex than usual, doing things others considered excessive or risky, and spending money in ways that caused trouble, among others. Each “yes” counts as one point, giving a symptom score between 0 and 13.
The second section is a single question asking whether several of those endorsed symptoms ever happened during the same time period. This is the co-occurrence filter. Saying “yes” to scattered symptoms that occurred years apart is different from reporting a cluster of them appearing together, which is a hallmark of a manic or hypomanic episode.
The third section asks how much of a problem the endorsed symptoms caused in your life. The response options are typically “no problem,” “minor problem,” “moderate problem,” or “serious problem.” Only “moderate” or “serious” counts toward a positive screen under the standard algorithm.
The Standard Scoring Algorithm Step by Step
To score the MDQ using the original method, you work through the three sections in sequence. First, tally the number of “yes” responses in the 13-item checklist. If that number is seven or higher, move to the co-occurrence question. If the person answered “yes” to co-occurrence, check the impairment question. If the impairment rating is “moderate” or “serious,” the screen is positive. If any of those three gates is not passed, the screen is negative.
This gated approach means the MDQ is deliberately conservative. A person could endorse 10 of the 13 symptom items but still screen negative if they said the symptoms never happened at the same time, or if they rated the resulting problems as only minor. That conservatism is intentional: the original developers wanted the screen to be fairly specific, meaning it should not flag too many people who do not actually have bipolar disorder.1PubMed Central. The Mood Disorder Questionnaire: A Simple, Patient-Rated Screening Instrument for Bipolar Disorder The tradeoff is that the screen misses some genuine cases, a problem explored in more detail below.
How Well the Standard Cutoff Actually Performs
The standard seven-symptom, three-gate algorithm does not perform identically across every setting. A large review pooling results from multiple studies found that across all settings the MDQ had an overall sensitivity of about 61% and specificity of about 88%.2PubMed. Screening for bipolar disorder with the Mood Disorders Questionnaire: a review In plain terms, the standard scoring catches roughly six out of ten people who truly have bipolar disorder and correctly clears about nine out of ten who do not.
Those numbers shift depending on where the MDQ is used. In psychiatric outpatient clinics, sensitivity tends to be higher because the patients there are more likely to have pronounced symptoms. In the general population, sensitivity drops substantially while specificity climbs: fewer true cases get caught, but the people who do screen positive are more likely to genuinely have bipolar disorder.2PubMed. Screening for bipolar disorder with the Mood Disorders Questionnaire: a review
One study in a psychiatric outpatient sample found that sensitivity for bipolar I was around 69%, but for bipolar II and related conditions it dropped to 30%.3PubMed. Sensitivity and specificity of the Mood Disorder Questionnaire for detecting bipolar disorder That gap matters because bipolar II, which involves hypomania rather than full-blown mania, is precisely the condition most often missed and most often misdiagnosed as plain depression. The MDQ’s symptom items are weighted toward classic manic experiences, and people with milder hypomanic episodes may not endorse enough of them to reach the seven-item threshold.4PubMed Central. Critical Overview of Screening Tools for Detecting Bipolar Disorders
The Impairment Question as a Source of False Negatives
A substantial share of false negatives, where people with bipolar disorder screen negative, trace back to the impairment question rather than the symptom count. In the same psychiatric outpatient study mentioned above, patients’ low ratings of how much trouble their symptoms caused explained nearly half of all false negatives.3PubMed. Sensitivity and specificity of the Mood Disorder Questionnaire for detecting bipolar disorder People who have lived with hypomania for years may not consider their elevated energy or reduced need for sleep to be problems. Some even view those experiences as productive periods. The impairment gate can therefore filter out exactly the people the screen is supposed to catch.
This is one reason researchers have explored whether dropping the supplementary questions (the co-occurrence and impairment items) and relying on symptom count alone improves detection. In a UK validation study, using a higher symptom threshold of nine or more endorsed items without the supplementary questions yielded sensitivity around 90% for both bipolar I and bipolar II, with specificity also around 90%.5PubMed. Validation of the Mood Disorder Questionnaire for screening for bipolar disorder in a UK sample Similarly, a study of pregnant and postpartum women found that a threshold of seven or more symptoms, without the supplementary questions, produced 89% sensitivity and 84% specificity.6The Journal of Clinical Psychiatry. Sensitivity and Specificity of the Mood Disorder Questionnaire as a Screening Tool for Bipolar Disorder During Pregnancy and the Postpartum Period
When Lower Symptom Thresholds Make Sense
The standard seven-symptom cutoff is not sacred. Several studies have found that lowering the threshold to five or even six symptoms, sometimes with or without the supplementary gates, improves the MDQ’s ability to catch cases without wrecking specificity.
In a study of primary care patients already diagnosed with depression, a cutoff of five symptoms yielded sensitivity of 91% and specificity of 67%.7PubMed Central. Screening for Bipolar Disorder Symptoms in Depressed Primary Care Attenders: Comparison between Mood Disorder Questionnaire and Hypomania Checklist (HCL-32) Among a genetically high-risk Anabaptist population, a five-symptom method similarly increased diagnostic sensitivity with little loss of specificity, and lowering the threshold below five offered diminishing returns.8PubMed. Validity of the Mood Disorder Questionnaire (MDQ) as a screening tool for bipolar spectrum disorders in anabaptist populations A Thai validation study also found that five positive items on the symptom checklist was the optimal threshold, yielding sensitivity around 77% and specificity around 73%.9PubMed Central. Development and validation of a screening instrument for bipolar spectrum disorder: The Mood Disorder Questionnaire Thai version
A Chinese-language validation kept the seven-symptom cutoff but assessed it without the impairment gate and found sensitivity of 73% and specificity of 88% for detecting bipolar disorder in a psychiatric population.10PubMed. Validation of the Chinese version of the Mood Disorder Questionnaire in a psychiatric population in Hong Kong A Rwandese adaptation recommended a cutoff of six symptom items and dropped two of the original items entirely to improve performance in that context.11PubMed. Adaption and validation of the Rwandese version of the Mood Disorder Questionnaire for the screening of bipolar disorder
The pattern across these studies is consistent: the standard algorithm prioritizes specificity at the expense of sensitivity, and when the goal is to catch more true cases rather than minimize false alarms, relaxing the symptom threshold or removing the supplementary gates helps. Which approach is better depends on the clinical context. In a busy primary care office where an undetected bipolar diagnosis could mean years of ineffective antidepressant treatment, casting a wider net may be worth the extra false positives. In a research study trying to recruit a clean bipolar sample, the stricter standard method is preferable.
What Causes False Positives
A positive MDQ screen does not mean you have bipolar disorder. This is a screening instrument, not a diagnostic one, and a meaningful number of people who screen positive will not meet diagnostic criteria when evaluated further. Understanding what drives false positives helps explain why the MDQ is a starting point, not an endpoint.
Among non-bipolar patients who incorrectly screened positive, several conditions independently predicted a false positive result: current PTSD, borderline personality disorder, current or recent substance use disorder, and a history of childhood abuse.12PubMed. Factors associated with false positives in MDQ screening for bipolar disorder: Insight into the construct validity of the scale All of these conditions can produce symptoms that overlap with the MDQ’s item list. Irritability, impulsivity, sleep disturbance, risk-taking behavior, and rapid shifts in mood or energy are features of PTSD, borderline personality disorder, and substance misuse as well as bipolar disorder. The MDQ cannot distinguish the source of those symptoms.
One critical analysis highlighted that when the MDQ cutoff was lowered to five symptoms (to achieve at least 90% sensitivity) without the impairment gate, specificity fell to about 61% and the positive predictive value dropped to roughly 22%. That means nearly four out of five people flagged by the screen at that threshold did not have bipolar disorder.13PubMed. Are screening scales for bipolar disorder good enough to be used in clinical practice? The authors of that study argued this low positive predictive value raises serious questions about using the MDQ as a standalone decision tool in clinical practice. That concern is worth taking seriously: a positive screen should always lead to a full clinical evaluation rather than a presumptive diagnosis.
Self-Report vs. Clinician Review of the Same Answers
An interesting wrinkle in MDQ scoring is that patients and clinicians often disagree about what the patient’s own answers mean. In a study of inpatients with mood symptoms and substance misuse, 56% of patients scored positive on the self-rated MDQ, compared to only 30% when clinicians reviewed the same patients’ responses.14Journal of Clinical Psychiatry. Clinician-rated versus self-rated screening for bipolar disorder among inpatients with mood symptoms and substance misuse
The disagreement was not spread evenly across items. Patients and clinicians had the hardest time agreeing on irritability, racing thoughts, and distractibility, while they agreed most on items like excessive spending, increased goal-directed activity, and hypersexuality.14Journal of Clinical Psychiatry. Clinician-rated versus self-rated screening for bipolar disorder among inpatients with mood symptoms and substance misuse The items with low agreement are the vague, subjective ones. Almost anyone can recall a time they felt irritable or distracted. The items with high agreement tend to involve concrete, observable behaviors. This suggests that when you are filling out the MDQ yourself, the more behavioral items, like spending sprees, sexual indiscretions, or sudden bursts of goal-directed work, carry more diagnostic weight than items describing internal states that everyone experiences to some degree.
Scoring the MDQ for Adolescents
The original MDQ was developed and validated in adults. When researchers adapted it for adolescents, they tested three versions: one where the adolescent reported their own symptoms, one where the adolescent guessed how a teacher or friend would describe them, and one completed by a parent about the adolescent’s symptoms. The parent-completed version performed best. A symptom threshold of five or more items on the parent version yielded sensitivity of 72% and specificity of 81%, which was superior to the self-report and the “how would others describe me” versions.15PubMed. Validation of the Mood Disorder Questionnaire for bipolar disorders in adolescents
The lower threshold makes sense given that adolescents have had fewer years to accumulate the full range of manic experiences the 13 items describe, and their insight into their own behavior tends to be less developed than that of adults. If you are scoring an adolescent version, the five-item parent-report threshold is a more validated starting point than the adult seven-item standard.
How the MDQ Compares to the Hypomania Checklist
The MDQ is not the only screening tool for bipolar disorder, and the most commonly compared alternative is the Hypomania Checklist (HCL-32), a 32-item instrument focused specifically on hypomanic symptoms. A meta-analysis comparing the two found that at study-defined cutoffs, the HCL-32 had sensitivity of about 82% and specificity of about 57%, while the MDQ had sensitivity of about 80% and specificity of about 70%.16PubMed. Comparison of the screening ability between the 32-item Hypomania Checklist (HCL-32) and the Mood Disorder Questionnaire (MDQ) for bipolar disorder: A meta-analysis and systematic review In other words, the two tools catch a similar proportion of true cases, but the MDQ produces fewer false positives.
A head-to-head study in medicated patients with major depressive disorder found a more dramatic split: the HCL-32 achieved 100% sensitivity but only 46% specificity at its optimal cutoff, while the MDQ hit 71% sensitivity and 77% specificity.17PubMed. Comparison of the validity of the Chinese versions of the Hypomania Symptom Checklist-32 (HCL-32) and Mood Disorder Questionnaire (MDQ) for the detection of bipolar disorder in medicated patients with major depressive disorder The HCL-32 flagged every bipolar case but also flagged nearly half the non-bipolar patients. In a primary care study comparing the two instruments among depressed patients, the MDQ again showed better overall accuracy and was quicker to administer.7PubMed Central. Screening for Bipolar Disorder Symptoms in Depressed Primary Care Attenders: Comparison between Mood Disorder Questionnaire and Hypomania Checklist (HCL-32)
In practical terms, neither tool is clearly superior. The HCL-32 may catch more bipolar II cases because its items are more tuned to hypomania, but at the cost of more false alarms. The MDQ is shorter, faster, and more specific, which makes it a better fit for settings where clinician time is limited and false positives are costly. Some clinicians use both and treat agreement between the two as a stronger signal than either alone.
Whether All 13 Items Are Equally Important
The standard scoring method treats every endorsed item as worth one point, but not all items contribute equally to distinguishing bipolar disorder from other conditions. A machine-learning analysis found that the best classification performance was achieved when all 13 items were included, meaning no single item or subset outperformed the full set.18PubMed Central. Exploring Informative Items for Bipolar Disorder Classification Using Machine Learning With Anger Coping Styles in Combination With the Mood Disorder Questionnaire and Bipolar Spectrum Diagnostic Scale This confirms that the MDQ works as a package deal rather than as a collection of individually diagnostic symptoms.
That said, the items are not interchangeable in how reliably people endorse them. As noted earlier, items describing internal experiences like irritability, distractibility, and racing thoughts tend to have low agreement between self-report and clinician assessment, while items about observable behaviors show much higher agreement. When interpreting a borderline score of six or seven, a clinician might reasonably give more weight to endorsed behavioral items than to endorsed subjective ones, even though the formal scoring method counts them equally.
Automated Scoring and Electronic Records
Increasingly, the MDQ is administered electronically through patient portals or tablets in waiting rooms rather than on paper. Several clinical software platforms now score the MDQ automatically and feed the result into the patient’s electronic health record. The scoring logic is the same: tally the symptom count, check co-occurrence, check impairment level, flag positive or negative. Automated scoring eliminates arithmetic errors and ensures that the supplementary gates are applied consistently, which is a genuine advantage over hand-scoring in a busy clinic where a staff member might tally symptom items but forget to check the impairment question.
In primary care and trauma-exposed populations, where the MDQ’s sensitivity with standard scoring hovers around 62% and its positive predictive value can be as low as 17%, the risk of over-relying on automated results is real.19PubMed. Diagnosing bipolar disorder in trauma exposed primary care patients An automated negative result could reassure a provider when it should not. The screen’s role is the same regardless of how it is administered: it narrows the field. A negative screen does not rule out bipolar disorder, especially bipolar II, and a positive screen does not confirm it.