Singing depends on a chain of biological systems working in tight coordination: the lungs push air upward, a pair of small tissue folds in the larynx vibrate to create a raw buzzing tone, and the throat, mouth, and nasal passages reshape that buzz into the rich, recognizable sound we hear as a voice. No single organ “sings.” The ability emerges from the interplay of respiratory muscles, laryngeal tissues, resonating cavities, and a brain that can learn to coordinate all of them with remarkable precision. What makes the human voice unusual is less any one anatomical feature and more the degree of voluntary control the brain exerts over these parts.
How Vocal Folds Produce Sound
Your vocal folds (sometimes called vocal cords, though they are not cords) are two small shelves of layered tissue that sit inside the larynx, stretched horizontally from front to back. When you breathe normally, they stay apart. When you speak or sing, muscles pull them together so that the gap between them, the glottis, narrows. Air pressure from the lungs builds beneath the closed folds until it forces them apart, releasing a puff of air. The folds then snap back together, and the cycle repeats, sometimes hundreds of times per second.
This process is called myoelastic-aerodynamic vibration. The “myoelastic” part refers to the muscular and elastic properties of the folds themselves; the “aerodynamic” part refers to the role of airflow and pressure in driving them open and pulling them shut. The folds do not vibrate because a nerve fires for each cycle. Instead, the brain sets the tension and position of the folds, and physics does the rest.
The vibration is not a simple flapping. A wave travels through the soft tissue of each fold from bottom to top, creating a rolling motion along the surface. Measurements of this “mucosal wave” in laryngeal tissue show propagation speeds ranging from about 0.5 to 2.0 meters per second, depending on the frequency of vibration.1The Journal of the Acoustical Society of America. Measurements of mucosal wave propagation and vertical phase difference in vocal fold vibration This wave is critical: it allows the lower edge of each fold to open before the upper edge, creating a phase difference that efficiently converts airflow into acoustic energy.
To change pitch, singers adjust the tension and length of the vocal folds. Imaging studies of singers performing across two octaves show that the folds first elongate and then increase in tension as pitch rises, changing their cross-sectional shape from thick to thin and back again over the range.2PubMed. Changes in Vocal Fold Morphology During Singing Over Two Octaves Longer, thinner, tighter folds vibrate faster, producing a higher pitch, much like tightening a guitar string.
Air Pressure and Volume
Loudness in singing comes primarily from pushing more air pressure beneath the vocal folds. When subglottal pressure increases, the folds are blown apart more forcefully with each cycle, creating a stronger sound wave. Computational models of voice production confirm that increasing subglottal pressure is the main driver of vocal intensity, though it also tends to push the pitch slightly upward and can introduce more noise into the signal.3PubMed Central. Cause-effect relationship between vocal fold physiology and voice production in a three-dimensional phonation model Measurements in trained singers show that louder singing involves higher subglottal pressure, greater airflow, and a larger amplitude of each glottal pulse.4PubMed. Glottal Adduction and Subglottal Pressure in Singing
The lungs and the muscles that control them are not passive bellows. Trained classical singers manage their breathing differently from untrained people. Research comparing the two groups during singing found that classical singers use proportionally less rib cage contribution to lung volume and demonstrate a distinctive coordination pattern in which the abdominal wall leads the rib cage during exhalation, rather than the two moving in lockstep. This produced lower average airflow and more controlled volume changes during singing of a standard piece.5PLoS ONE. Breathing and Singing: Objective Characterization of Breathing Patterns in Classical Singers In plain terms, trained singers are parceling out their air supply more efficiently, giving them longer phrases and steadier sound.
How the Throat and Mouth Shape the Sound
The raw sound produced at the vocal folds is a harsh, buzzy tone. It only becomes a recognizable vowel, a warm singing tone, or a piercing operatic call after it passes through the vocal tract: the throat, mouth, tongue, and sometimes the nasal cavity. These spaces act as acoustic filters, amplifying some frequencies and dampening others. The frequencies that get boosted are called formants, and their positions depend on the shape of the tract at any given moment. Moving the tongue, jaw, lips, or soft palate changes which frequencies ring out and which get suppressed.
A common assumption is that when singers adjust these resonating spaces, they are somehow improving the vibration of the vocal folds themselves. Computational modeling suggests this is mostly wrong. The improvements singers achieve by shaping their vocal tract come primarily from changes in the tract’s acoustic response, not from changes in the sound source at the glottis.6PubMed Central. The influence of source-filter interaction on the voice source in a three-dimensional computational model of voice production In other words, the throat and mouth are doing most of the heavy lifting in shaping vocal quality.
One of the most studied examples of this shaping is the “singer’s formant,” a prominent peak in the frequency spectrum near 3 kHz found in the voices of classical operatic singers. It is produced by a clustering of several higher formants, creating a concentrated band of energy that allows a solo voice to cut through the sound of a full orchestra.7PubMed. Level and center frequency of the singer’s formant The singer’s formant region also exhibits greater directivity than lower frequencies, meaning it projects forward more strongly. Singers appear to control this directivity by adjusting how much spectral energy they place in that frequency band.8PubMed. Long-term horizontal vocal directivity of opera singers: effects of singing projection and acoustic environment
Throat Singing and Extreme Vocal Tract Shaping
Tuvan throat singing demonstrates just how far vocal tract shaping can go. In biphonic throat singing, a performer produces two distinct pitches simultaneously: a low drone from the vocal folds and a high, flute-like overtone isolated by extremely precise positioning of the tongue. Research using MRI and acoustic analysis shows that the singer creates two narrow constrictions in the vocal tract. The retroflex position of the tongue tip produces a constriction at about 14 centimeters from the lips and opens a sublingual space beneath, and it is the degree of constriction at these two locations that allows the singer to select and amplify a single high overtone.9PubMed Central. Overtone focusing in biphonic tuvan throat singing
MRI data from overtone singers confirms that for low enhanced overtone frequencies, the tongue tip is raised and strongly retracted, while for higher overtones, the tongue forms a longer but less retracted constriction. The front cavity of the vocal tract acts like a tunable resonator, its frequency set by the tongue tip position and lip opening.10PubMed. Voice source, formant frequencies and vocal tract shape in overtone singing. A case study Throat singing is not a different kind of voice production; it uses the same vocal folds and the same airstream. The extraordinary part is the vocal tract manipulation, which takes an ability everyone’s anatomy possesses and pushes it to a limit most people never approach.
The Brain’s Role in Pitch Control
Vocal fold tension, tongue position, jaw opening, breathing rhythm: all of these must be coordinated in real time. The brain region most directly responsible for controlling vocal pitch is the dorsal laryngeal motor cortex, a patch of cortical tissue that selectively encodes produced pitch. Neural recordings in humans show that this area controls short pitch accents used for prosodic emphasis in speech and also encodes pitch during singing. Direct electrical stimulation of this region evokes involuntary vocalization, confirming its causal role.11PubMed Central. The Control of Vocal Pitch in Human Laryngeal Motor Cortex
Interestingly, brain imaging of singers producing different pitches found that the peak activation sites in the larynx motor cortex were remarkably consistent across pitch levels within the same person, with identical or immediately adjacent activation peaks for different notes. The brain does not appear to maintain a spatial map where high notes live in one spot and low notes in another.12PubMed Central. How does human motor cortex regulate vocal pitch in singers? Instead, the same neural territory handles the full pitch range, likely using patterns of neural firing rather than physical location to encode different targets.
Humans are unusual among primates in having this kind of direct cortical control over laryngeal muscles. A review of comparative neuroscience describes changes to the location, structure, function, and connectivity of the larynx motor cortex in humans compared with other primates, including features that underpin our capacity for voluntary vocal learning.13PubMed Central. The origins of the vocal brain in humans Most non-human primates have only indirect cortical connections to their laryngeal muscles, which is one reason they can produce innate calls but cannot learn new vocal patterns the way we can.
What Training Actually Changes in the Brain and Body
Everyone with a functioning larynx can sing after a fashion. The difference between a trained and untrained singer is partly muscular and partly neural. A scoping review of research on singing skill learning found consistent evidence that singing expertise is associated with enhanced integration between auditory and motor brain systems, reduced reliance on external auditory feedback, and greater stability of vocal output when sensory conditions are disrupted. These changes involve functional and structural adaptations across a distributed network including auditory, motor, somatosensory, and subcortical brain regions.14Journal of Voice. Neuroplasticity and Neural Adaptations in Singing Voice Skill Learning: A Scoping Review
One of the most telling demonstrations involves artificially shifting the pitch of a singer’s own voice as they hear it through headphones. When this happens, untrained singers immediately adjust their pitch to compensate for what they hear. Trained singers, however, compensate less during the initial shift. After the altered feedback is removed, singers show a lingering aftereffect: their pitch remains slightly shifted, suggesting they had begun to update an internal model rather than simply chasing the auditory signal in real time. This aftereffect even transfers when the singer switches to a different note, indicating the internal model is abstract rather than tied to one motor pattern.15PubMed Central. Auditory-motor mapping for pitch control in singers and nonsingers In short, trained singers rely more on an internal sense of where their voice should be, while untrained singers rely more on what they hear.
Longitudinal work tracking singing students over the course of their education found that auditory feedback’s contribution to pitch accuracy did not improve with training, but the contribution of kinesthetic feedback (the bodily sense of laryngeal position and tension) did improve in certain singing tasks.16PubMed. Effects of a professional solo singer education on auditory and kinesthetic feedback–a longitudinal study of singers’ pitch control This helps explain why experienced singers can maintain good intonation even in noisy environments where they can barely hear themselves.
Physical Exercise for the Voice
The vocal folds contain muscle tissue, and like any muscle, it responds to use and disuse. Research in zebra finches, one of the best-studied vocal learning species, shows that daily vocal exercise is necessary to first develop and then maintain peak performance of the vocal muscles. When male birds were experimentally prevented from singing, both the physiology and performance of their vocal muscles declined within days. Females preferred the songs of vocally exercised males, suggesting that vocal output acts as an honest signal of recent exercise investment.17Nature Communications. Daily vocal exercise is necessary for peak performance singing in a songbird While the leap from finch to human is a large one, the underlying principle is consistent with what vocal coaches have long observed: singers who stop practicing for extended periods lose muscular conditioning quickly, even if their technique knowledge remains intact.
What Determines Your Vocal Range
Vocal fold length is a good predictor of average pitch but a poor predictor of total range. A cross-species study of vocal capacity found that what determines how wide a range an animal can achieve is not overall fold size but the properties of the fibrous layers within the folds. When vocal fold tissues develop a layered structure with fibers running lengthwise, the densest and stiffest layer governs the range. The stress-strain curve of this fibrous layer needs to be highly nonlinear to overcome the natural tendency for pitch to drop as folds are stretched longer.18PLoS Computational Biology. Predicting Achievable Fundamental Frequency Ranges in Vocalization Across Species Human vocal folds have a well-developed vocal ligament with exactly this kind of nonlinear stiffness, which is part of why we can cover a wider pitch range than our larynx size alone would predict.
Sex hormones reshape the larynx substantially during puberty and continue to influence the voice throughout life. Testosterone drives the growth of the male larynx at puberty, lengthening the vocal folds and dropping the voice about an octave. Female voices drop less but still change. Hormonal effects on the voice extend to the menstrual cycle, pregnancy, and aging.19PubMed Central. Effect of sex hormones on human voice physiology: from childhood to senescence Research on the register break between chest and falsetto voice suggests a sex-based pattern, with female voices showing smaller characteristic leap intervals and less individual diversity than male voices when transitioning between registers.20PubMed Central. Measurement of characteristic leap interval between chest and falsetto registers
Why Some People Cannot Carry a Tune
Roughly 4 percent of the population has congenital amusia, a hereditary condition marked by a specific deficit in processing musical pitch. A family aggregation study found that about 39 percent of first-degree relatives of people with amusia share the condition, compared with only 3 percent of relatives in control families.21The American Journal of Human Genetics. The Genetics of Congenital Amusia (Tone Deafness): A Family-Aggregation Study Brain imaging studies show that amusia involves reduced activation of the right inferior frontal gyrus and weakened connections between this region and the auditory cortex, rather than a problem in the auditory cortex itself.22PubMed. Functional MRI evidence of an abnormal neural network for pitch processing in congenital amusia
A surprising wrinkle: many people with congenital amusia can still make unconscious vocal pitch adjustments. When their vocal feedback is suddenly shifted in pitch through headphones, a majority of amusics still produce a corrective response, and nearly half respond with normal timing and magnitude. The size of this response was predicted by vocal pitch matching accuracy rather than the ability to consciously perceive small pitch changes.23PubMed. Vocal pitch shift in congenital amusia (pitch deafness) This points to a dual-route system in the brain: one route for conscious pitch perception and another, partly independent route for the fast, automatic adjustments that keep the voice on track. You can have a broken perception route and still retain some automatic vocal control.
How Voices Break Down
The same collision forces that make singing possible can also damage the vocal folds when the load is excessive or chronic. Vocal fold nodules, the small callous-like growths familiar to many professional singers and teachers, develop as a cumulative tissue response to repetitive impact. Each vibration cycle involves the two folds slamming together, and the collision force concentrates at the midpoint of the membranous fold. Over time, the nonlinear stiffening of the vocal ligament under chronic stress leads to localized tissue remodeling.24Journal of Voice. Biomechanical Mechanisms of Vocal Fold Nodule and Polyp Formation: A Review Nodules tend to be symmetric and bilateral, appearing at the same spot on both folds, because that midpoint is where the mechanical stress is greatest. Polyps, by contrast, often result from a single hemorrhagic event rather than gradual wear.
Why Humans Sing and Birds Do Too
Comparisons between human and bird vocal production reveal striking parallels. Songbirds generate sound in the syrinx, an organ located where the trachea branches into the two bronchi, rather than in a larynx. Despite this anatomical difference, the physical mechanism is fundamentally similar: both systems rely on tissue masses that are brought together and then set vibrating by airflow, with the flow rate modulated by the oscillation. Both systems feature layered tissue structures with complex elastic properties.25PubMed Central. Peripheral mechanisms for vocal production in birds – differences and similarities to human speech and singing
One key difference is that songbirds have two independently controlled sound sources, one on each side of the syrinx, allowing them to produce two independent sounds simultaneously or switch rapidly between the two. The songbird vocal apparatus is also adapted for much higher speed modulation than the human system, which makes sense given that temporal patterns and rapid frequency sweeps play a central role in birdsong communication. These similarities are a case of convergent evolution rather than shared ancestry: the last common ancestor of birds and mammals did not sing. The two lineages independently evolved layered vibrating tissues, voluntary neural control of those tissues, and a capacity for learned vocal patterns, arriving at the same functional solution through very different anatomical routes.
The human larynx also has an evolutionary story of its own. Comparative anatomy of primates shows that the descent of the hyoid bone within the neck, which creates the elongated pharyngeal cavity humans use for both speech and singing, occurred specifically during hominid evolution.26PubMed. Comparative morphology of the hyo-laryngeal complex in anthropoids: two steps in the evolution of the descent of the larynx A lower larynx gives us a longer resonating tube, which in turn gives us a wider range of formant frequencies and greater acoustic diversity. The trade-off is a shared airway and food passage that makes choking possible, a risk other primates largely avoid. The fact that evolution preserved this arrangement suggests that the vocal advantages it conferred were significant enough to outweigh the danger.