Hearing Different Words Than What Is Said: A Perception Test

Every healthy brain routinely hears words that were never spoken, and this is not a sign of anything wrong. Speech perception is a constructive process: your brain does not passively record incoming sound waves like a microphone but instead actively predicts, fills in, and sometimes overrides the raw acoustic signal. Several well-studied illusions demonstrate this vividly, and researchers use them as perception tests to explore how tightly hearing, vision, context, and expectation are woven together when you process language.

Your Brain Is Guessing, Not Just Listening

The idea that you hear exactly what hits your eardrums is one of the most intuitive and most wrong assumptions about human perception. In reality, your auditory cortex works in tandem with language centers and even motor regions associated with producing speech. When you hear someone talk, parts of your brain that control your own mouth and tongue fire as if simulating the movements needed to make those sounds. This simulation generates a prediction of what should come next, and that prediction shapes what you actually perceive. Brain imaging studies have shown that even during silent articulation, auditory cortex in the upper part of the temporal lobe lights up in both hemispheres, reflecting forward predictions about what speech should sound like.

This predictive machinery is why speech perception works so well in noisy rooms, on bad phone connections, and across wildly different accents. But it is also why your brain can confidently deliver a word to your conscious awareness that was never actually uttered. The perception tests that reveal this tendency are not party tricks; they expose the fundamental architecture of how you understand language.

The McGurk Effect and What Your Eyes Do to Your Ears

The most famous demonstration of hearing a word that was not said is the McGurk effect. In this illusion, you watch a video of someone mouthing one syllable while the audio track plays a different one. A common pairing uses a visual “ga” with an auditory “ba,” and most people confidently report hearing “da,” a sound that exists in neither the video nor the audio. The brain, faced with conflicting signals, does not pick one input and ignore the other. Instead it fuses them into a compromise that feels completely real.

What makes this illusion useful as a perception test is the enormous range of individual responses. In a study examining what drives these differences, researchers found that a person’s susceptibility to the McGurk effect was linked to their ability to read fine-grained mouth movements, but not to attention, processing speed, or working memory.1PLoS ONE. What accounts for individual differences in susceptibility to the McGurk effect? In other words, if you are naturally good at picking up visual speech cues, your brain gives those cues more weight when they clash with the audio, and the illusion hits harder. People who rely less on lip-reading tend to hear the actual audio more often.

A large cross-cultural study comparing native Mandarin Chinese speakers and native American English speakers found remarkably similar rates of the McGurk effect: about 48% and 44% respectively, a difference that was not statistically meaningful. Individual variation within each group was far larger than any difference between them, with some people experiencing the illusion on every trial and others never falling for it at all.2PubMed Central. Similar frequency of the McGurk effect in large samples of native Mandarin Chinese and American English speakers The takeaway is that this fusion of sight and sound is not culturally learned in any simple way. It appears to be a deep feature of how human brains handle speech.

Phonemic Restoration and the Brain’s Autocorrect

You do not need conflicting visual information for your brain to fabricate speech sounds. In the phonemic restoration effect, a single sound within a word is physically removed and replaced by a noise burst, like a cough or a tone. Listeners do not notice the gap. They hear the complete, intact word, and they can even point to where the noise occurred as separate from the word rather than inside it. The illusion is so strong that people will argue they heard the missing sound clearly.

This effect becomes more powerful when the surrounding sentence provides context about which word should appear. Research using brain recordings has shown that supportive sentence context affects speech processing at a surprisingly early stage, influencing the perception of individual speech sounds before higher-level meaning-making kicks in.3PubMed Central. The phonemic restoration effect reveals pre-N400 effect of supportive sentence context in speech perception Your brain essentially fills in the blank using what it expects should be there. When the context is strong, the fill-in is nearly irresistible.

This is not an error in any meaningful sense. It is a design feature. In natural environments, speech is constantly interrupted by ambient noise, and if your brain waited for every phoneme to arrive intact before assembling a word, conversation would be nearly impossible. Phonemic restoration keeps speech perception fluid and continuous. The cost is that sometimes the brain’s autocorrect inserts the wrong word entirely, especially when the acoustic signal is degraded or when you have strong expectations about what someone is about to say.

The Verbal Transformation Effect

A different kind of perceptual drift happens when you listen to a single word repeated on a loop. Within seconds to minutes, most people begin hearing the word change into something else, then something else again, cycling through a series of phantom words that are not in the recording. This is called the verbal transformation effect, and it has been studied since the mid-twentieth century.

In one experiment, 192 participants listened through headphones to repeating words presented at different spatial locations. They reliably reported independent illusory changes at each location, hearing different phantom words in each ear or at the center.4PubMed. The verbal transformation effect: auditory illusions as an index of lexical processing and homolog activation The illusions were not random; they tended to be real words or near-words in the listener’s language, suggesting that the brain’s lexical network is actively generating candidates and occasionally “winning” over the actual repeated stimulus.

You can test this yourself with any repeating audio clip online. The phenomenon is related to a broader category of perceptual satiation, where sustained exposure to an unchanging stimulus causes your neural representation of it to become unstable. It is a vivid reminder that what you hear at any given moment is not a fixed readout of the sound wave but a constantly updated best guess.

Why Context Changes What You Hear

Expectations do not just nudge perception at the margins; they can fundamentally determine what you perceive. This is especially clear in experiments involving ambiguous or degraded audio. Sine-wave speech, for instance, strips natural speech down to a few frequency-tracking tones that sound like electronic chirps. People who are told the sounds are random noise typically hear gibberish. But once told they are listening to speech and given an example, many suddenly hear clear words in the exact same recording. Brain imaging shows that this shift is accompanied by a dramatic increase in high-frequency neural activity and that frontal brain regions begin discriminating between specific words only after listeners understand that speech is present.5PubMed Central. Neural correlates of sine wave speech intelligibility in human frontal and temporal cortex The acoustic input did not change. The listener’s model of what to expect changed, and perception followed.

This learning effect can be remarkably durable. Classic studies on synthetic speech found that after about eight days of training with feedback, listeners went from roughly 20% accuracy to about 70% accuracy, and they retained this gain even six months later without additional exposure.6Frontiers. Speech perception as an active cognitive process What is striking is that the improvement generalized to words the listeners had never practiced on. They had not memorized individual words; they had restructured how their perceptual system decoded the signal.

The flip side is that priming can make you hear things that are not there. In a study on “electronic voice phenomena,” participants who were told they were listening to recordings with paranormal content reported hearing significantly more voices in ambiguous audio compared to participants who were simply told they were hearing degraded speech. Same audio, different framing, different perception.7PLOS ONE. Paranormal experiences, sensory-processing sensitivity, and the priming of pareidolia This is auditory pareidolia, the hearing equivalent of seeing faces in clouds, and it illustrates how powerful top-down expectations are in shaping what you perceive.

Misheard Lyrics and Mondegreens

Perhaps the most familiar everyday version of hearing different words is the mondegreen, a term for misheard song lyrics. The classic example involves Jimi Hendrix’s “Purple Haze,” where “‘Scuse me while I kiss the sky” becomes “‘Scuse me while I kiss this guy.” These mishearings are not random. They tend to produce plausible phrases, and once a particular interpretation locks in, it can be very hard to un-hear.

Brain imaging research has explored what happens neurally during these misperceptions. When people experience mondegreens or their cross-language equivalent (called soramimi, where lyrics in one language sound like phrases in another), the brain activates a bilateral network that includes middle temporal and inferior frontal areas as well as the anterior cingulate cortex and parts of the thalamus. Interestingly, anterior cingulate activity correlated with how amusing listeners found the misperceptions, linking the surprise of hearing unexpected words to the brain’s reward and conflict-monitoring systems.8PubMed Central. Neurobiology of knowledge and misperception of lyrics

Mondegreens thrive in conditions where the acoustic signal is ambiguous: fast tempos, unfamiliar accents, distorted instrumentation, or lyrics that are acoustically similar to alternative phrases. Once you know the correct lyrics, you usually hear them correctly, which again demonstrates the power of top-down knowledge in resolving ambiguous input. But for someone who has never seen the lyrics printed, the misheard version may persist for years.

How Aging Shifts the Balance

The balance between bottom-up acoustic processing and top-down contextual prediction changes as people age. Older adults tend to rely more heavily on sentence context to understand speech, compensating for declining auditory acuity. In a study comparing younger and older adults on speech recall tasks, the older group showed significantly greater use of semantic context on both immediate repeat and delayed recall tasks.9PubMed Central. The effect of aging on context use and reliance on context in speech: A behavioral experiment with Repeat–Recall Test

This heavier reliance on context is often helpful. It allows older adults to maintain functional comprehension even when their ears miss some of the fine acoustic detail. But it also means they are more vulnerable to hearing the wrong word when the context is misleading or when background noise introduces competing information. If you have ever noticed an older relative confidently responding to something nobody actually said, this mechanism is often the explanation. The brain filled in based on context, and the fill-in happened to be wrong.

Peripheral hearing loss adds another layer. When certain frequency ranges are attenuated, the brain receives a more ambiguous signal, which gives the predictive system more room to impose its own interpretation. Hearing aids help by restoring some of that missing acoustic detail, but they do not eliminate the perceptual filling-in process. The brain continues to construct its best guess from whatever signal it gets.

Audiovisual Integration and Autism

Because the McGurk effect depends on merging visual and auditory information, researchers have used it to study how multisensory integration works in people with different neurological profiles. In autism spectrum conditions, there has been considerable debate about whether audiovisual speech integration operates differently. One early study found no significant difference in the rate of fused percepts between children with autism and typically developing controls when audiovisual stimuli were presented simultaneously.10PubMed Central. Multisensory Speech Perception in Children with Autism Spectrum Disorders

However, a more recent meta-analysis pooling 18 studies with a combined sample of 952 participants found that autistic individuals did show reduced audiovisual integration compared to non-autistic peers, with a moderate-to-large effect size.11PubMed. Differences between autistic and non-autistic individuals in audiovisual speech integration: A systematic review and meta-analysis The discrepancy between smaller studies and the meta-analytic picture is a reminder that individual study results on perception can be noisy, and that the size and methodology of studies matters. The practical implication is that autistic individuals may rely less on lip movements to disambiguate speech, which could affect comprehension in visually rich but acoustically noisy environments differently than it does for neurotypical listeners.

White Noise, False Voices, and What It Does Not Mean

One of the more anxiety-inducing versions of hearing words that are not there involves white noise. Some people report hearing speech or fragments of words embedded in static. This has occasionally been proposed as a screening tool for psychosis risk, based on the logic that perceiving structure in random noise might reflect the same tendency that produces hallucinations.

But the evidence for that link is shaky. A study using a general-population twin sample found that perceiving speech in white noise was not significantly associated with either positive or negative schizotypy.12PubMed Central. White noise speech illusion and psychosis expression: An experimental investigation of psychosis liability In other words, hearing phantom words in static appears to be a normal perceptual phenomenon, not a red flag for mental illness. The brain is constantly searching for patterns in auditory input, and sometimes it finds speech-like patterns where none exist. The priming study on paranormal recordings mentioned earlier shows that this tendency can be amplified by suggestion and expectation in perfectly healthy people.

If you have ever heard faint words in a running shower, a fan, or a washing machine, you are experiencing the same process. The auditory system is biased toward detecting speech because speech has been the single most important sound category in human evolutionary history. This bias means occasional false positives, and those false positives are not pathological.

Noisy Environments and the Cocktail Party Problem

Real-world speech perception is rarely as clean as a laboratory headphone experiment. In any environment with multiple talkers, your brain has to solve what researchers call the cocktail party problem: picking out one voice from a mix of competing speech and background noise. This task relies on grouping sounds by primitive features like spatial location and the fundamental pitch of each speaker’s voice, then selectively attending to one stream while suppressing others.13PubMed Central. The cocktail-party problem revisited: early processing and selection of multi-talker speech

When this separation is imperfect, fragments from a competing voice can intrude into the attended stream. Your brain may splice those fragments into the sentence you think you are following, producing a coherent but wrong utterance. This is one of the most common sources of everyday “hearing different words”: not a single-speaker illusion, but a multi-speaker blending error. The effect worsens in reverberant rooms, at louder ambient noise levels, and when the competing voice has a similar pitch to the target speaker.

These real-world factors also interact with the aging effects discussed earlier. An older adult with even mild hearing loss in a noisy restaurant faces a compounding problem: reduced acoustic input, greater reliance on context, and more opportunities for cross-talk between competing voices. It is not a failure of intelligence or attention; it is a perfectly predictable consequence of how perception works under degraded conditions.

Where Your Brain Puts It All Together

The brain does not have a single “speech perception center.” Instead, different parts of the superior temporal sulcus handle different aspects of the task. Research using brain imaging has mapped out a gradient along this region: anterior and middle portions respond preferentially to auditory speech signals, while posterior portions respond more to visual speech cues like lip movements. Mid-temporal areas show the strongest preference for speech over non-speech sounds in general.14PubMed Central. Auditory, Visual and Audiovisual Speech Processing Streams in Superior Temporal Sulcus This distributed layout means that audiovisual integration is not a one-step process where sight and sound merge in a single spot. It is a coordinated handoff across a strip of cortex, with different zones weighting different inputs.

This architecture has evolutionary roots. Both human and non-human primate brains contain pathways specialized for processing dynamic social signals like facial expressions and mouth movements and integrating them with vocal cues.15PubMed Central. Faces and Voices Processing in Human and Primate Brains: Rhythmic and Multimodal Mechanisms Underlying the Evolution and Development of Speech The system that merges what you see on someone’s face with what you hear from their mouth predates language. Humans repurposed it for speech, which is partly why visual information has such a powerful grip on auditory perception and why illusions like the McGurk effect are so hard to override even when you know the trick.

Deepfakes and the Limits of Perceptual Trust

The same perceptual biases that produce speech illusions create vulnerabilities in the age of synthetic media. Deepfake videos manipulate both the visual and auditory channels, and human perception is not well equipped to catch the mismatch. In a study presenting 110 participants with a mix of real and deepfake videos, every AI detection model tested outperformed human judges on the same set of clips.16arXiv. Unmasking Illusions: Understanding Human Perception of Audiovisual Deepfakes Both native and non-native English speakers struggled to distinguish real from fake.

This is a direct consequence of the constructive nature of speech perception. Your brain is designed to smooth over minor inconsistencies between sight and sound, to fill in degraded audio with plausible words, and to trust coherent-looking lip movements. Those are the exact features deepfake technology exploits. When the lip movements roughly match the audio, your perceptual system accepts the package as genuine because that is what it has been doing automatically your entire life. Detecting deepfakes may ultimately require tools that bypass human perception entirely, precisely because perception was built to believe.