The House-Brackmann Score: A Facial Nerve Grading System

The House-Brackmann (HB) score is a six-point clinical scale that rates facial nerve function from Grade I (normal) to Grade VI (total paralysis), and it has been the most widely used system for this purpose since it was adopted by the American Academy of Otolaryngology–Head and Neck Surgery in 1985. Its appeal lies in its simplicity: a clinician watches a patient make a few facial expressions, assigns a single grade, and that number becomes shorthand for how much the nerve is working. But the score’s simplicity is also its main limitation, and a growing body of research shows that different clinicians looking at the same face often disagree on the grade by a meaningful margin.

What Each Grade Means

The system divides facial nerve function into six tiers. Grade I means completely normal movement with no weakness at all. Grade II is mild dysfunction, where there is slight weakness noticeable only on close inspection, and the eye can still close completely. Grade III represents moderate dysfunction, with an obvious but not disfiguring difference between the two sides of the face and complete eye closure with effort. Grade IV describes moderately severe dysfunction, with obvious weakness and incomplete eye closure. Grade V is severe dysfunction, where barely perceptible movement is present. Grade VI is total paralysis, with no movement at all.1PubMed Central. Recovery rate and prognostic factors of peripheral facial palsy treated with integrative medicine treatment: a retrospective study

In clinical practice, the grades also function as recovery benchmarks. Reaching Grade I means the patient has achieved complete recovery. Grade II is generally considered a good functional outcome, meaning there is some residual asymmetry that most people in daily conversation would not notice. Grades III and IV occupy a middle zone where movement is clearly impaired but meaningful voluntary motion remains. Grades V and VI indicate severe or total loss of function, and these patients face the greatest risk of secondary complications, particularly around the eye.

Where the Score Gets Used

The HB score shows up across a surprisingly wide range of clinical scenarios, from emergency rooms to operating theaters to rehabilitation clinics. Its most frequent appearance is in Bell’s palsy, the sudden onset facial paralysis that accounts for the majority of facial nerve problems. Clinicians grade a patient at the first visit, then re-grade at follow-up appointments to track whether the nerve is recovering. In a large study of Bell’s palsy outcomes, patients who presented with initial HB Grades III or IV had a significantly higher rate of favorable recovery at six months compared to those who started at Grades V or VI, with roughly 83% versus 68% reaching a good outcome.2PubMed Central. Association between Initial Severity of Facial Weakness and Outcomes of Bell’s Palsy

In neurosurgery, the score is the standard language for reporting facial nerve outcomes after tumor removal. Vestibular schwannomas (tumors on the balance nerve that often press on the facial nerve) are a common example. In a series of patients with larger tumors, good long-term facial nerve function, defined as HB Grade I or II, was achieved in over 90% of cases regardless of whether the resection was total, near-total, or subtotal.3Neurosurgery. Facial Nerve Preservation Surgery for Koos Grade 3 and 4 Vestibular Schwannomas That 90% figure is what patients and surgeons discuss when weighing the risks of surgery, and the HB score is the metric that makes the conversation possible.

Trauma is another major context. When a blow to the head damages the temporal bone near the ear, the facial nerve can be injured in the process. In a study of patients with temporal bone injuries who developed facial palsy, the mean initial HB grade was about 2.8, and patients recovered to a mean of roughly 1.6 over an average follow-up of about 25 days.4PubMed Central. Clinical Features of Fracture versus Concussion of the Temporal Bone after Head Trauma That relatively modest starting grade reflects the fact that traumatic facial palsy is often incomplete at onset, which generally carries a better prognosis than complete paralysis.

What the Initial Grade Tells You About Recovery

One of the most practically important questions patients ask is: will my face get better? The initial HB grade is one of the strongest clinical predictors. A higher initial grade (meaning worse paralysis) is associated with slower and less complete recovery. Research into Bell’s palsy prognosis confirms that the initial HB grade, along with time to treatment onset and whether synkinesis develops, are among the strongest predictors of final outcome.5PubMed Central. Predicting recovery in Bell’s palsy: the impact of early rehabilitation and prognostic indicators

Electrical nerve testing can add prognostic information that the HB grade alone cannot provide. Electroneurography (ENoG) measures how much of the nerve’s electrical signal has been lost, and electromyography (EMG) detects the type of nerve injury at the muscle level. In one study, follow-up EMG showed the best prognostic accuracy at 97%, outperforming ENoG, which reached about 73%.6PubMed. Prognostic value of electroneurography and electromyography in facial palsy When the HB score was compared directly to ENoG findings over time in Bell’s palsy patients, both told a similar story by the follow-up period, though the HB grades initially suggested slightly milder impairment than the electrical tests did.7PubMed. House-Brackmann and Yanagihara grading scores in relation to electroneurographic results in the time course of Bell’s palsy The practical takeaway is that the HB grade gives a useful first impression, but electrical testing adds a layer of objectivity that helps when big decisions, like whether to consider surgery, are on the table.

The Reliability Problem

The biggest criticism of the HB system is that different clinicians do not consistently agree on which grade to assign. The system asks raters to make a holistic judgment about the whole face, which involves subjective decisions about whether a movement is “slight,” “obvious,” or “barely perceptible.” Those adjectives leave room for interpretation. In a study comparing several grading systems across six raters evaluating the same patients twice, the HB system produced only moderate interrater agreement, with a Fleiss’ kappa around 0.4 to 0.5.8PubMed Central. Comparison of the Reliability of the House–Brackmann, Facial Nerve Grading System 2.0, and Sunnybrook Facial Grading System for the Evaluation of Patients with Peripheral Facial Paralysis Another study of the HB system found that while weighted kappa indicated substantial overall reliability, the exact agreement between trained observers was only about 44%.9PubMed. Reliability of the “Sydney,” “Sunnybrook,” and “House Brackmann” facial grading systems to assess voluntary movement and synkinesis after facial nerve paralysis

Interestingly, the disagreement is not evenly distributed across the face. When researchers broke down reliability by facial region, agreement was strongest for the midface and for the global score (kappa around 0.5), but noticeably weaker for the forehead and eye regions (kappa around 0.3). Clinicians with more experience tended to agree more closely with each other.10PubMed. Significance and reliability of the House-Brackmann grading system for regional facial nerve function This matters because the eye region is often the most clinically important area to evaluate accurately, since incomplete eye closure can lead to corneal damage.

These agreement numbers might sound abstract, but think about what they mean in practice. If two doctors see the same patient and one assigns Grade III while the other assigns Grade IV, that single grade of difference can change treatment decisions: whether to start steroids more aggressively, whether to refer for surgery, whether to pursue eye protection measures. A system that produces this level of disagreement in nearly half of assessments has real consequences for patient care.

How It Compares to Other Grading Systems

The HB scale is not the only option. The Sunnybrook Facial Grading System uses a continuous scoring approach that separately rates resting symmetry, voluntary movement across five facial regions, and synkinesis (involuntary movement that appears when a patient tries to move one part of the face and another part moves too). The result is a composite score from 0 to 100 rather than a single grade. The Facial Nerve Grading System 2.0 (FNGS 2.0) was developed as a direct update to the original HB system, adding regional scoring to address some of its limitations.

In head-to-head reliability comparisons, the Sunnybrook system and FNGS 2.0 consistently outperform the original HB scale. Both showed excellent interrater agreement (ICC above 0.95) compared to the HB’s moderate agreement in the same study.8PubMed Central. Comparison of the Reliability of the House–Brackmann, Facial Nerve Grading System 2.0, and Sunnybrook Facial Grading System for the Evaluation of Patients with Peripheral Facial Paralysis The trade-off is speed. The HB took about one minute per patient on average, the FNGS 2.0 about a minute and a half, and the Sunnybrook system about two and a half minutes. In a busy clinic, that difference matters.

A large analysis of over 5,000 paired facial gradings showed that while the two systems correlated well overall, Sunnybrook scores spread widely within each HB grade. For example, patients assigned HB Grade III had Sunnybrook composite scores ranging from roughly 43 to 62 when looking at the middle 50% of results, and HB Grade IV patients ranged from about 26 to 43.11PubMed. Sunnybrook and House-Brackmann systems in 5397 facial gradings That spread means two patients who look meaningfully different on the Sunnybrook scale can land in the same HB grade, confirming that the six-point scale collapses clinically relevant differences. Correlation was weakest at the initial visit and strongest at later follow-ups, suggesting the systems agree more once the nerve has settled into a stable state.

Despite its limitations, the HB system persists as the default for a practical reason: nearly every published study on facial nerve outcomes over the past four decades reports results using it. Switching to a different system would make it harder to compare new research against the existing literature. This backward compatibility keeps the HB scale firmly in place even as evidence mounts that its competitors are more precise.

Eye Complications and What the HB Grade Misses

Corneal damage is one of the most serious consequences of facial paralysis. When the orbicularis oculi muscle around the eye does not close properly, the cornea is exposed to drying, debris, and infection. You might expect that a worse HB grade would reliably predict which patients will develop these eye problems, but the relationship turns out to be surprisingly weak. Research found no HB grade cutoff that was both sensitive and specific enough to screen for corneal complications. In other words, the numeric HB grade alone is not a useful guide for identifying patients at risk of corneal damage from facial palsy.12PubMed. The House-Brackmann system and assessment of corneal risk in facial nerve palsy

That said, the HB score is not useless for eye assessment. When combined with other measurements such as the degree of lagophthalmos (gap between the eyelids when trying to close) and loss of corneal sensation, the HB scale does become a clinically significant predictor of whether a patient will need protective eye surgery.13PubMed. Ocular factors predicting requirement for corneal protective oculoplastic surgery The lesson is that no single number captures the full picture. Eye risk assessment needs its own focused evaluation beyond whatever the HB grade says.

Does the Score Match How Patients Actually Feel?

A clinician looking at a face and a patient living behind that face often have different views of how things are going. Quality of life instruments like the Facial Clinimetric Evaluation (FaCE) scale and the Facial Disability Index (FDI) capture the patient’s own experience of their facial paralysis, including emotional distress, social withdrawal, and difficulty eating or drinking. A systematic review and meta-analysis pooled correlations between these patient-reported measures and clinician-graded scales. The HB score correlated moderately with the FaCE total score, with a pooled correlation of about 0.53, and with FDI physical function at about 0.47.14JAMA Otolaryngology–Head & Neck Surgery. Associations Between Clinician-Graded Facial Function and Patient-Reported Quality of Life in Adults With Peripheral Facial Palsy: A Systematic Review and Meta-analysis

Those correlations are meaningful but far from strong. The weakest link was with social function: the pooled correlation between the HB score and FDI social function was only about 0.17. That is barely a connection at all, which makes sense if you think about it. Social functioning depends on factors like how self-conscious the patient feels, how others react to their face, and whether synkinesis makes them reluctant to smile in public. None of that is captured by a clinician watching someone raise their eyebrows in an exam room. This gap is why many facial nerve centers now use patient-reported questionnaires alongside clinician grading rather than relying on the HB score alone.

Using the Scale in Children

Grading facial nerve function in children presents unique challenges. Young children may not follow instructions well, may cry during examination, and may be difficult to assess at rest. In a study of 174 children with Bell’s palsy, researchers tested a modified parent-administered version of the HB scale against the standard clinician-administered version. The overall agreement was good, with a mean intraclass correlation of 0.88 across all time points. Agreement was weakest at the initial visit (ICC of 0.53) and strongest at the one-month follow-up (ICC of 0.88), suggesting that parents are less confident in grading when the paralysis is new and most distressing.15PubMed Central. Agreement of Clinician-Administered and Modified Parent-Administered House-Brackmann Scales in Children with Bell’s Palsy

A secondary analysis from the same pediatric trial compared HB scores to Sunnybrook scores in children and found high agreement overall (ICC of 0.92), though agreement was again weak at baseline (ICC of 0.37). Clinicians rated the HB scale as significantly easier to administer than the Sunnybrook, with a median ease score of 3 versus 7 on a difficulty scale.16PubMed. Agreement Between House-Brackmann and Sunnybrook Facial Nerve Grading Systems in Bell’s Palsy in Children: Secondary Analysis of a Randomized, Placebo-Controlled Multicenter Trial For pediatric settings, where speed and simplicity are especially valuable, the HB scale’s ease of use is a genuine advantage even if its precision is lower than alternatives.

Machine Learning and the Push Toward Automation

One emerging approach to the reliability problem is to take clinicians out of the grading loop entirely, or at least give them an objective second opinion. The auto-eFACE is a machine-learning tool that analyzes photographs and video of a patient’s face and produces standardized scores. In testing, the automated system successfully distinguished normal faces from those with facial palsy, though it found minor asymmetries in “normal” faces that clinicians tend to overlook, producing slightly lower scores for healthy controls than the perfect 100 that clinicians assigned. For patients with flaccid paralysis and severe synkinesis, the automated tool tended to score facial symmetry slightly higher than clinicians did.17Plastic and Reconstructive Surgery. The Auto-eFACE: Machine Learning–Enhanced Program Yields Automated Facial Palsy Assessment Tool

The promise of automated systems is not just convenience but consistency. A computer analyzing the same photograph twice will always produce the same score, eliminating the rater-to-rater variability that plagues subjective scales. Whether automated grading can fully replace clinical judgment is another question, since a machine may miss contextual factors that an experienced clinician catches intuitively, like whether a patient’s resting asymmetry is longstanding rather than new. But as these tools mature, they could serve as a reliable baseline that clinicians use alongside their own assessment, much the way imaging has become an expected complement to the physical exam in other areas of medicine.

Surgical Reanimation and Treatment Thresholds

When facial nerve recovery stalls at a poor HB grade despite conservative management, surgical reanimation becomes a consideration. The HB score serves as both the trigger for these discussions and the metric for evaluating outcomes. One surgical technique, the jump interpositional graft hypoglossal-facial anastomosis, achieved HB Grade III or better in about 83% of properly selected patients, which represents a meaningful improvement from pre-surgical paralysis.18Laryngoscope. Facial reanimation with jump interpositional graft hypoglossal facial anastomosis and hypoglossal facial anastomosis: evolution in management of facial paralysis Notice that the benchmark for surgical success here is Grade III, not Grade I. For patients starting from total paralysis, regaining moderate function with obvious but manageable asymmetry is considered a win. That calibration of expectations, using the same grading language the patient has been hearing since diagnosis, is one of the HB score’s underappreciated strengths: it gives patients and surgeons a shared vocabulary for discussing realistic goals.

Three grading systems evaluated the same patients with facial palsy using standardized facial expressions, and all three showed statistically significant changes after three months, confirming that whichever scale you use, the trajectory of recovery is captured.19PubMed Central. Comparison of 3 Grading Systems (House-Brackmann, Sunnybrook, Sydney) for the Assessment of Facial Nerve Paralysis and Prediction of Neural Recovery The debate, ultimately, is not whether any of these scales work at all, but which one loses the least information in the process of reducing a complex, three-dimensional, emotionally loaded human face to a number. The HB score loses more than its competitors, but it has been losing it consistently in the same way for forty years, and that consistency has a value of its own.