The speech recognition threshold, or SRT, is the softest level at which a person can correctly repeat back spoken words about half the time. Audiologists measure it in decibels and use it as one of the foundational pieces of a hearing evaluation, alongside the pure-tone audiogram. The test sounds simple enough, but the details of how it works, what it reveals, and where it falls short tell you a lot about the complexities of human hearing.
What the Test Actually Measures
During an SRT test, you sit in a sound-treated booth wearing headphones, and the audiologist plays a series of two-syllable words at decreasing volume levels. The goal is to find the exact intensity, measured in decibels hearing level (dB HL), at which you can correctly repeat about 50% of the words. That 50% correct point is your speech recognition threshold. It tells the clinician the faintest level at which speech is barely intelligible to you, not comfortably heard, just barely understood.
The SRT is tested for each ear separately, and results are compared against your pure-tone audiogram. If your hearing is relatively straightforward, such as a flat or gently sloping hearing loss, the SRT should land close to your pure-tone average, the average of your hearing thresholds at key speech frequencies. Research consistently shows a strong correlation between these two measures, and clinicians routinely use SRT to validate whether the pure-tone results are reliable.1PubMed Central. Prediction of pure tone thresholds using the speech reception threshold and age in elderly individuals with hearing loss
Why Spondaic Words Are Used
The words used during an SRT test are not random vocabulary. They are spondees, two-syllable words where both syllables carry roughly equal stress, like “baseball,” “hotdog,” or “toothbrush.” The reason is practical: spondees have a steep psychometric function, meaning that as the volume increases even slightly, the jump from not understanding the word to recognizing it happens over a narrow range. That steep transition makes it easier to pinpoint the exact threshold, because the response does not gradually drift upward over a wide range of intensities.
Not all spondees perform equally well. Research in the 1980s found that out of the 36 spondaic words commonly used in clinics, only about 15 were truly homogeneous, meaning they all became intelligible at roughly the same sound level and their intelligibility grew at a similar rate. The good news: SRT measurements using just those 15 well-matched words were just as accurate as measurements using the full list of 36.2PubMed. Thresholds and psychometric functions of the individual spondaic words Clinicians have also studied whether patients need to be familiarized with the word list before testing. One study found that certain spondees were correctly identified even without prior familiarization, and the researchers built a shortened list of words whose recognition rates were minimally affected by whether the patient had heard them beforehand.3PubMed. A spondee list for determining speech reception threshold without prior familiarization
The takeaway is that audiologists are not just picking words at random. Decades of psychometric work have gone into selecting and refining spondee lists so that SRT results are as precise and reproducible as possible.
How SRT Relates to the Pure-Tone Audiogram
If you have ever had a hearing test, you probably remember raising your hand or pressing a button when you heard beeps at different pitches. Those results form your pure-tone audiogram. The pure-tone average, typically the average of your thresholds at 500, 1000, and 2000 Hz, should closely match your SRT. When it does, the audiologist has greater confidence that both tests are valid.
When the two numbers diverge significantly, it raises a flag. In straightforward hearing loss, the gap between the pure-tone average and the SRT is usually small, just a few decibels. But when there is a large mismatch, particularly when the SRT is much better than the pure-tone average, it can suggest that the pure-tone results are not reflecting true hearing ability. One study of cooperative patients with hearing loss found the average gap between a three-frequency pure-tone average and SRT was only about 2 dB, but in cases where patients were not responding honestly, the gap ballooned dramatically.4PubMed. Identification of pseudohypacusis using speech recognition thresholds That same study found a two-frequency pure-tone average was actually more effective at catching exaggerated hearing loss than the traditional three-frequency average.
Detecting Exaggerated Hearing Loss
The SRT–pure-tone comparison has long been one of the audiologist’s sharpest tools for identifying what is sometimes called pseudohypacusis or non-organic hearing loss, where a patient reports worse hearing than they actually have. This can happen in medicolegal cases, workers’ compensation claims, or situations where someone has a conscious or unconscious incentive to appear more hearing-impaired than they are.
A hypothesis put forward in research on functional hearing loss explains why the SRT tends to be significantly better than the pure-tone average in these cases. The theory suggests that people faking or exaggerating hearing loss use a loudness criterion when deciding whether to respond: they wait until the stimulus sounds “loud enough.” Because of differences in how pure tones and speech are calibrated, spondee words and pure tones that are at the same sound pressure level seem equally loud, but the calibration offsets used in audiometry mean the speech stimulus effectively gets a head start. The result is that a person responding based on perceived loudness rather than actual detectability will inconsistently set their response thresholds across the two tests.5PubMed Central. Pure tone-spondee threshold relationships in functional hearing loss: a hypothesis
How the Test Procedure Works Step by Step
Several standardized protocols exist for measuring SRT. The American Speech-Language-Hearing Association (ASHA) published guidelines in 1979 and updated them in 1988, and research has compared the two approaches.6PubMed. A comparison of American Speech-Language Hearing Association guidelines for obtaining speech-recognition thresholds In general, the procedure works like this: the audiologist starts by presenting spondees at a level well above where the patient is expected to hear them. The level is then decreased in steps, often 5 or 10 dB at a time at first, then in smaller 2 dB steps as the threshold is approached. The patient repeats each word, and the audiologist tracks correct responses. The threshold is defined as the lowest level at which the patient gets about 50% of words correct.
There is some flexibility in step size and starting level, and different protocols handle these details slightly differently. However, the core logic is the same across versions: bracket the threshold from above, narrow in with smaller steps, and identify the 50% point.
Testing Children
Standard spondee word lists were designed for adults, so they do not always work well for young children who may not have the vocabulary to repeat words like “staircase” or “doormat.” Audiologists have developed several adaptations. One common approach for younger kids uses picture-pointing tasks: the child sees a board with images and points to the picture matching the word they heard, rather than repeating it verbally.
A Spanish-language pediatric SRT test, for example, was developed using this picture-pointing method for Spanish-speaking children. The resulting threshold typically fell within about 2 to 12 dB of the child’s pure-tone average, confirming that the test was measuring what it was supposed to.7PubMed. Spanish Pediatric Speech Recognition Threshold Test Another approach uses digit pairs rather than spondees. A study of 30 children aged 5 to 8 measured SRT using both paired digits and standard pediatric word stimuli, with different step sizes, to see which combinations worked best for young ears.8PubMed. Digit speech recognition threshold (SRT) in children with normal hearing ages 5-8 years
More recently, bilingual considerations have entered the picture. A test called the Children’s English and Spanish Speech Recognition Test (ChEgSS) was developed and validated for children as young as four, providing normative data for both Spanish/English bilingual and English monolingual children across a wide age range.9PubMed Central. Development of the Children’s English and Spanish Speech Recognition Test: Psychometric Properties, Feasibility, Reliability, and Normative Data
Why Language and Accent Matter
SRT testing is inherently tied to the language the patient speaks. Standard clinical word lists in the United States were developed for native English speakers. If you are a non-native speaker, unfamiliar spondees can trip you up not because of hearing loss but because of vocabulary. Research has shown that digit pairs, rather than traditional spondee words, produce more accurate SRT results for non-native English speakers. Because digits like “three-nine” or “five-two” are among the most universally familiar English words, they sidestep the vocabulary problem while still measuring hearing sensitivity for speech.10American Journal of Audiology. Digit Speech Recognition Thresholds (SRT) for Non-Native Speakers of English
This extends to speech-in-noise testing as well. When researchers compared native and non-native German speakers on several speech recognition tasks, they found that non-native listeners could achieve native-like thresholds only on the simplest, most linguistically basic material, namely digit triplets. For more complex sentence-based tests, non-native speakers needed the speech signal to be 3 to 6 dB louder relative to the background noise to achieve the same 50% recognition level as native speakers.11PubMed. How much does language proficiency by non-native listeners influence speech audiometric tests in noise? A 3 to 6 dB difference may sound small, but in signal-to-noise terms it is meaningful enough to change a clinical result from pass to fail. Clinicians testing multilingual patients need to choose materials carefully or risk confusing language proficiency with hearing ability.
This is also why multiple countries and languages have developed their own validated spondee lists. An Urdu spondee word list, for instance, went through rigorous psychometric evaluation with native speakers, testing each word’s slope and threshold characteristics before finalizing a clinically usable list.12PubMed. Psychometric Evaluation of Digitally Recorded Urdu Spondee Word List for Speech Reception Threshold Testing
SRT in Quiet Versus SRT in Noise
The standard SRT test is performed in quiet, a silent booth with no competing sounds. But real life is rarely quiet. You listen to conversations in restaurants, on busy streets, and in rooms with music playing. That mismatch has pushed researchers toward speech-in-noise versions of the SRT, where words or sentences are presented against a background of noise. The threshold is then expressed as a signal-to-noise ratio rather than an absolute decibel level: it tells you how much louder speech needs to be than the noise for you to understand half of what is said.
Peripheral hearing loss, meaning damage to the sensory cells or auditory nerve fibers in the inner ear, contributes to elevated SRTs both in quiet and in noise.13Hearing Research. Speech recognition in noise and presbycusis: relations to possible neural mechanisms But some people have hearing thresholds that look normal on a standard audiogram yet still struggle to follow speech when there is background noise. This is sometimes called “hidden hearing loss,” and it suggests that the quiet-booth SRT can miss real-world hearing difficulties. Speech-in-noise SRT testing captures a dimension of hearing that the standard test does not.
When Hearing Aids Are in the Picture
SRT measurements also play a role in hearing aid fitting and evaluation, though with some important caveats. Adaptive SRT tests in the clinic typically present speech at signal-to-noise ratios that are worse than what people actually encounter in daily life. Research has shown that at higher, more realistic signal-to-noise ratios, these SRT tests become insensitive to differences in hearing aid processing that may genuinely benefit the listener.14Ear and Hearing. Using Speech Recall in Hearing Aid Fitting and Outcome Evaluation Under Ecological Test Conditions In other words, the SRT test might show no difference between hearing aid settings even though one setting is clearly better in everyday listening. This is an active area of research, with clinicians exploring supplementary measures like speech recall tasks to better evaluate hearing aid benefit under real-world conditions.
Cochlear Implant Referrals
For people with more severe hearing loss, SRT testing feeds into decisions about cochlear implant candidacy. Current practice guidelines suggest that hearing professionals should consider referring adults for a cochlear implant evaluation when certain audiometric criteria are met. A “60/60 guideline” has been proposed, encouraging referral when specific hearing benchmarks are reached, with the hope that clearer referral criteria will improve access to cochlear implants for people who could benefit from them.15Otology & Neurotology. Development of a 60/60 Guideline for Referring Adults for a Traditional Cochlear Implant Candidacy Evaluation Speech recognition performance, including SRT and word-recognition scores, is a major part of the evaluation process once a referral is made.
Auditory Neuropathy and Unusual Patterns
In most types of hearing loss, the SRT and the audiogram tell a consistent story. But in a condition called auditory neuropathy spectrum disorder (ANSD), the picture gets messier. People with ANSD may have relatively normal inner ear function, with intact outer hair cells, but disrupted signaling along the auditory nerve. Their audiograms can look surprisingly good, yet their ability to understand speech, especially in noise, may be severely impaired.
Research on children with ANSD who received cochlear implants found that their SRT scores in quiet and in noise were similar to those of implanted children with standard sensorineural hearing loss.16PubMed. Rate of neural recovery in implanted children with auditory neuropathy spectrum disorder However, the underlying perceptual profiles are quite different. Children with ANSD tend to have normal frequency resolution but disrupted temporal processing, the ability to track rapid changes in sound over time. This temporal disruption was strongly correlated with their speech perception performance, meaning that the timing dimension of hearing, rather than pitch or loudness sensitivity, was the key bottleneck.17Ear and Hearing. Perceptual Characterization of Children with Auditory Neuropathy Standard SRT testing can miss this distinction because it was designed to measure sensitivity, not temporal precision.
Cognitive Factors in Speech Recognition
Hearing is not entirely an ear problem. Your brain does an enormous amount of processing to turn acoustic signals into meaningful language, and cognitive resources like working memory and attention play a role, particularly when listening conditions are tough. Research on listeners with normal hearing found that both age and working memory capacity had significant effects on speech recognition in noise. Working memory span was the single most important cognitive variable predicting performance, and the interaction between age and working memory was particularly strong for sentence-level material: older adults with lower working memory showed worse speech-in-noise scores, while those with high working memory performed more like younger listeners.18Ear and Hearing. Effects of Age and Working Memory Capacity on Speech Recognition Performance in Noise Among Listeners With Normal Hearing
Neuroimaging research adds another layer. A study of older adults found that the relationship between frontal lobe cortical thickness and speech-in-noise thresholds depended on the listener’s working memory capacity. Older adults with both thicker frontal cortex and greater working memory capacity achieved the best speech-in-noise results, suggesting the brain’s structural resources and cognitive capacity work together.19Neuropsychologia. Interacting effects of frontal lobe neuroanatomy and working memory capacity to older listeners’ speech recognition in noise That said, the picture is not perfectly tidy. Another study of elderly listeners found that the expected links between cognition and speech recognition emerged only under specific spatial-listening conditions, and no significant link with working memory was found in the hearing-impaired group at all.20Frontiers in Psychology. Exploring the Link Between Cognitive Abilities and Speech Recognition in the Elderly Under Different Listening Conditions The role of cognition in speech understanding is real but far from fully mapped.
Remote and Automated Testing
Teleaudiology has grown rapidly, and SRT testing has followed. Researchers have developed automated SRT tests that use forced-choice responses and computerized scoring, aiming to produce results that closely match traditional in-clinic pure-tone thresholds.21PubMed. Automated Forced-Choice Tests of Speech Recognition Self-administered audiometry applications have also been validated: one equivalence study found that remote testing of both frequency discrimination and speech recognition thresholds produced results equivalent to those obtained in a sound-treated booth.22PubMed. Validation of a Self-Administered Audiometry Application: An Equivalence Study
An automated speech-in-noise screening test designed for remote use showed high test-retest reliability and was accurate in identifying ears with hearing thresholds above 25 dB HL.23PubMed. An Automated Speech-in-Noise Test for Remote Testing: Development and Preliminary Evaluation These tools could be particularly valuable for screening large populations or reaching people in areas without easy access to an audiology clinic. The technology is not a full substitute for a face-to-face evaluation, where the audiologist can observe the patient’s behavior and adjust testing in real time, but it is narrowing the gap enough to be clinically useful for initial screening and follow-up monitoring.
Bone Conduction SRT
Most SRT testing is done through headphones or speakers, delivering sound by air conduction. But SRT can also be measured via bone conduction, where a vibrator placed behind the ear sends sound directly through the skull to the inner ear. Bone conduction SRT testing helps audiologists distinguish between types of hearing loss: if your air conduction SRT is worse than your bone conduction SRT, some of the problem is in the outer or middle ear rather than the inner ear. Early clinical studies found a correlation of about 0.75 between binaural SRT and the better ear’s pure-tone average when both were measured by air conduction.24JAMA Otolaryngology–Head & Neck Surgery. Speech Audiometry by Bone Conduction Bone conduction SRT adds diagnostic specificity by letting clinicians separate sensory from conductive components of hearing loss without relying solely on pure-tone data.