What Are Vocalizations? A Biological Explanation

Vocalizations are sounds that animals actively produce using specialized structures in their respiratory tracts, driven by airflow and controlled by the brain. They are distinct from other biological sounds like heartbeats or footsteps because they require dedicated anatomy and neural coordination to generate signals that carry information. Virtually all birds and mammals can vocalize, along with many amphibians, reptiles, and fish, though the organs involved and the complexity of the sounds vary enormously across species. What makes vocalization especially interesting biologically is that it sits at the intersection of anatomy, neuroscience, genetics, and ecology, with each layer adding nuance to how and why animals make the sounds they do.

How the Vocal Apparatus Produces Sound

At its core, a vocalization is an aerodynamic event. In mammals, the sound source is the larynx, a cartilage-and-muscle structure in the throat containing the vocal folds (often called “vocal cords,” though they are folds of tissue, not cords). When air from the lungs is pushed upward, it sets the vocal folds vibrating. The classic explanation of this process, known as the myoelastic-aerodynamic theory, originally held that rising air pressure forces the folds apart and falling pressure lets them snap shut. That picture turned out to be incomplete. In-vivo measurements showed that the key driver of sustained vibration is not simple pressure swings but rather vertical phase differences in the tissue itself: the lower edge of each fold opens slightly before the upper edge, creating a wave-like ripple along the fold’s surface. This “mucosal wave” makes the fold’s shape more convergent when opening and more divergent when closing, efficiently transferring aerodynamic energy into vibration cycle after cycle.1PubMed. Integrative Insights into the Myoelastic-Aerodynamic Theory and Acoustics of Phonation. Scientific Tribute to Donald G. Miller In different voice registers the details shift. During falsetto-like phonation, where the tissue moves more uniformly, the inertia of the air column passing through the glottis becomes the main asymmetric force that keeps vibration going.2PubMed. Comments on the myoelastic – aerodynamic theory of phonation

Once the vocal folds generate a raw buzzing tone, it passes through the vocal tract, the open spaces of the throat, mouth, and nasal cavities, which act as acoustic filters. Certain frequencies are amplified by the resonant properties of these cavities, and others are dampened. This “source-filter” framework explains why two people (or two animals) with similarly sized vocal folds can still sound different: the shape and length of the tract above the folds sculpt the final sound. Computational models show that extreme constrictions at various points along the tract, at the false vocal folds, the pharynx, or the lips, reduce overall airflow amplitude, while narrowing near the larynx entrance can boost certain high-frequency harmonics under specific vocal fold conditions.3PubMed Central. The influence of source-filter interaction on the voice source in a three-dimensional computational model of voice production

Birds took a different evolutionary path. Instead of a larynx, they produce sound with a syrinx, an organ located where the trachea splits into the two bronchi, deep inside the chest. This placement turns out to be acoustically advantageous. Physical models comparing the two positions found that a sound source at the syringeal position required lower air pressure to start vibrating and produced louder sound across most tracheal lengths. The efficiency gap was largest for tracheal lengths in the range that many birds actually have.4PLOS Biology. The evolution of the syrinx: An acoustic theory Because the syrinx sits at a branching point, some birds can independently control airflow through each bronchus, allowing them to produce two different pitches simultaneously, something no mammalian larynx can do.

The Brain’s Role in Vocal Control

Producing a vocalization is not just a matter of blowing air past some tissue. It requires precise coordination of dozens of muscles controlling the diaphragm, ribcage, larynx, tongue, lips, and jaw, all firing in the right sequence and at the right time. The motor neurons that directly drive these muscles are scattered across the brainstem and spinal cord, from the trigeminal motor nucleus in the pons to nuclei in the medulla and ventral horn of the cervical, thoracic, and lumbar spinal cord. A sprawling network in the brainstem coordinates these neuron pools, integrating sensory feedback from the lungs, larynx, and mouth to keep the system running smoothly.5PubMed. Neural pathways underlying vocal control

But this coordinating network does not simply switch on by itself. It needs a “go” signal from a region in the midbrain called the periaqueductal gray (PAG). Damage to the PAG in mammals causes mutism, even though the muscles and brainstem circuits are perfectly intact. The PAG acts as a gate: it determines whether a vocalization happens at all. Above the PAG, higher brain structures add layers of control. The anterior cingulate cortex and nearby frontal regions govern whether you choose to start or suppress a vocalization. The motor cortex, via its connections to the brainstem, gives you voluntary control over the acoustic details of what you say, allowing you to shape pitch, timing, and articulation.5PubMed. Neural pathways underlying vocal control This layered architecture means that a pain shriek can bypass the cortex entirely, erupting as a reflexive brainstem response, while a carefully articulated sentence engages virtually the whole chain from cortex to diaphragm.

Comparative work across primates, rodents, and songbirds reveals both deeply conserved circuits and convergent solutions. Songbirds and humans independently evolved specialized forebrain pathways for learned vocalization, even though they are separated by hundreds of millions of years of evolution. Rodents and non-human primates, meanwhile, rely more heavily on the ancient brainstem-PAG system for their calls, with more limited cortical involvement.6PubMed Central. The neurobiology of innate, volitional and learned vocalizations in mammals and birds

Innate Calls Versus Learned Vocalizations

One of the most fundamental distinctions in the biology of vocalization is between sounds that are innate and those that are learned. An innate call is one an animal produces without ever hearing it from another individual. A newborn kitten’s cry, a chicken’s alarm call, a frog’s mating chorus: these emerge from genetically programmed neural circuits without requiring a tutor. A learned vocalization, by contrast, requires the animal to hear a model, memorize it, and then practice until its own output matches the template. Human speech is the most obvious example, but vocal learning also occurs in songbirds, parrots, hummingbirds, bats, cetaceans, pinnipeds, and elephants.

The parallels between songbird and human vocal learning are striking. Both require auditory experience during an early sensitive period to form accurate memories of adult sounds. Both progress from variable, messy immature vocalizations, babbling in human infants and “subsong” in young birds, to stable, crystallized patterns through practice.7PubMed Central. Birds and babies: Ontogeny of vocal learning Auditory feedback is critical at every stage. Studies in songbirds show that disrupting a bird’s ability to hear itself sing causes its song to gradually deteriorate, and deafening a young bird before it has finished learning prevents song crystallization altogether.8PubMed Central. The role of auditory feedback in vocal learning and maintenance Interestingly, even species that do not learn their calls still adjust them. Marmosets, which lack true vocal learning, shift aspects of their calls in response to changing acoustic conditions, pointing to rapid sensory-motor interactions that maintain vocal flexibility even without learning.8PubMed Central. The role of auditory feedback in vocal learning and maintenance

Despite how much vocal learning tells us about language, the research has been heavily skewed toward birds. A survey of published studies over 25 years found that roughly 84% of original research articles on vocal learning focused on birds, while bats, pinnipeds, cetaceans, and elephants together accounted for only about 8%.9Current Opinion in Behavioral Sciences. Vocal learning: a language-relevant trait in need of a broad cross-species approach That imbalance means our understanding of mammalian vocal learning outside of humans is still thin.

Body Size, Vocal Fold Length, and Pitch

There is a reason a mouse squeaks and a lion roars: larger animals tend to produce lower-pitched vocalizations. This relationship, known as acoustic allometry, exists because bigger bodies tend to house longer vocal folds, and longer folds vibrate at lower fundamental frequencies, much as a longer guitar string produces a deeper note. Among primates, vocal fold length turns out to be a far better predictor of minimum fundamental frequency than body size alone, explaining about 81% of the variation across species compared to roughly 40-50% for body mass.10Scientific Reports. Acoustic allometry revisited: morphological determinants of fundamental frequency in primate vocal production This matters because it means an animal’s voice carries honest information about its body. A rival hearing a deep call can reasonably infer that the caller is large, which is useful information when deciding whether to fight or flee.

The relationship is not perfect, though. Some species have evolved exaggerated vocal anatomy that lets them sound bigger than they are, or unusual mechanisms that decouple pitch from body size entirely. Mice, for instance, produce ultrasonic vocalizations well above 20 kHz, far higher than their body size would predict from simple fold vibration. Recent computational work suggests that mice generate these calls not by vibrating vocal folds in the conventional way but through a hole-edge mechanism involving airflow past a sharp tissue edge inside the larynx. A small internal cavity called the ventral pouch appears to play a key role in shaping the diversity of mouse ultrasonic call types.11PubMed. Contribution of the ventral pouch in the production of mouse ultrasonic vocalizations

How Hormones Shape the Voice

Sex hormones, particularly testosterone, exert powerful effects on vocal behavior and vocal anatomy. In many species, males vocalize more during the breeding season, and this seasonal surge in singing is closely tied to rising testosterone levels. But testosterone does more than just flip a motivational switch. In canaries, researchers found that testosterone implanted in a specific brain region (the medial preoptic nucleus) increased how often a castrated male sang but did not improve the acoustic quality of the song. Only when testosterone was allowed to circulate broadly through the brain did song structure become more complex and stereotyped.12PubMed Central. Differential effects of global versus local testosterone on singing behavior and its underlying neural substrate In other words, the hormone controls both the desire to vocalize and the neural refinement of the vocalization itself, through separate brain pathways.

Testosterone can also alter the physical properties of the sound. Zebra finches given testosterone implants showed drops in the frequency of song elements over several weeks, and the frequency change persisted even after the implants were removed.13PubMed. Testosterone implants alter the frequency range of zebra finch songs In wild black redstarts, blocking testosterone and estradiol during territorial encounters altered structural song measures like trill rate and frequency, suggesting that these hormones underpin the ability to modify songs in aggressive contexts.14PLOS ONE. Testosterone Affects Song Modulation during Simulated Territorial Intrusions in Male Black Redstarts (Phoenicurus ochruros) Researchers have proposed that this hormonal modulation of acoustic structure may be a general mechanism across vertebrates, not limited to birds.

Genetics and the FOXP2 Story

One of the clearest genetic links to vocal ability involves a gene called FOXP2. In humans, a mutation in this gene was discovered in a family (the “KE family”) in which affected members had severe difficulties with speech and language. FOXP2 is not a “language gene” in any simple sense; it encodes a transcription factor that regulates the activity of many other genes. But research in mice carrying the same mutation found in the KE family revealed something concrete about what goes wrong. The mutation disrupts intracellular “protein motors” in the striatum, a brain region critical for learning and executing motor sequences. Specifically, it causes an abnormal buildup of a protein called dynactin1, which impairs how neurons transport signaling molecules, grow dendrites, and fire electrically, alongside measurable deficits in vocalization. Knocking down the excess dynactin1 in these mice rescued both the cellular problems and the vocal deficits.15PubMed Central. Speech- and language-linked FOXP2 mutation targets protein motors in striatal neurons

FOXP2 is highly conserved across mammals and birds, and versions of it have been implicated in song learning in songbirds and echolocation call development in bats. Its story illustrates a broader point: vocalization depends on motor learning circuits in the brain, and the genes that build and maintain those circuits can have cascading effects on an animal’s ability to produce and refine sounds.

Evolutionary Origins of Acoustic Communication

How far back does vocalization go? A large comparative study mapping acoustic communication across the vertebrate family tree found evidence that it has ancient, shared roots. The African lungfish, a living relative of the earliest land vertebrates, can produce and perceive sounds both in water and in air, hinting that the basic capacity for acoustic signaling was present before vertebrates even left the water.16Nature Communications. Common evolutionary origin of acoustic communication in choanate vertebrates At the neural level, the evidence is even more compelling. A highly conserved pattern of hindbrain and spinal cord circuitry links the vocal systems of fish to those of all major lineages of vocal tetrapods, including frogs, reptiles, birds, and mammals. The proposal is that the neural compartment responsible for acoustic communication was already present in ancient fishes and was inherited, not independently invented, by their descendants.17PubMed Central. Evolutionary origins for social vocalization in a vertebrate hindbrain-spinal compartment

The vocal organs themselves, however, did evolve independently multiple times. The larynx of mammals and the syrinx of birds are not the same structure modified from a common ancestor; they arose separately. What was inherited was the neural blueprint for generating and controlling rhythmic vocal motor patterns. The specific tissue that vibrates is secondary to the brain circuitry that coordinates it.

When Sound Is Not a Vocalization

Not every sound an animal produces with its body counts as a vocalization. Biologists draw a line between vocalizations, which involve the vocal organ and airflow, and “sonations,” which are mechanical sounds produced by other body parts. Some of the most vivid examples come from birds. The Anna’s hummingbird produces a loud chirp during its courtship dive, but high-speed video revealed that the sound comes from its tail feathers, not its syrinx. The outermost tail feathers flutter at high speed in the airstream, generating the chirp mechanically.18PubMed Central. The Anna’s hummingbird chirps with its tail: a new mechanism of sonation in birds Streamertail hummingbirds produce a distinctive flight hum from modified wing feathers that generate tones in the 700-900 Hz range.19PubMed Central. Fluttering wing feathers produce the flight sounds of male streamertail hummingbirds

Perhaps the most remarkable case is the Club-winged Manakin, a small tropical bird whose males produce a sustained, tonal “ting” sound at about 1,500 Hz during courtship by rapidly vibrating specialized wing feathers against each other. The modified feather shafts are hypertrophied and show unusually high resonant tuning, with quality factors (a measure of how sharply tuned a resonator is) far exceeding those typical for biological structures.20PubMed Central. Resonating feathers produce courtship song These examples matter because they show that the pressure to communicate acoustically can drive the evolution of sound-producing structures well beyond the vocal organ. The biological need to be heard is not limited to the throat.

Singing in a Noisy World

Vocalizations do not exist in a vacuum. They have to travel through an environment that can distort, absorb, or mask them. One long-standing idea, the acoustic adaptation hypothesis, predicts that animals living in dense vegetation should produce lower-frequency, longer calls (because low frequencies travel better through foliage), while open-habitat species should use higher, more rapidly repeated signals. It is an elegant idea, but a recent meta-analysis found no overall support for it across terrestrial vertebrates. Neither within-species nor among-species comparisons showed a consistent effect of vegetation structure on call frequency or timing.21PubMed. Meta-analysis of the acoustic adaptation hypothesis reveals no support for the effect of vegetation structure on acoustic signalling across terrestrial vertebrates Some individual studies do find local support, such as a South American frog whose southern populations produced calls that degraded less in their local habitat compared to calls from other populations, while northern populations showed the opposite pattern.22PubMed Central. The acoustic adaptation hypothesis in a widely distributed South American frog: Southernmost signals propagate better But the grand picture is messier than the hypothesis predicts.

A more robust finding involves anthropogenic noise. Traffic, construction, and urban hum create a blanket of low-frequency sound that can drown out animal vocalizations. Multiple studies have documented birds shifting their songs to higher frequencies in noisy environments. Among tropical birds, eight out of nine species studied sang at higher dominant frequencies in forest fragments near urban areas with heavy traffic noise.23PubMed. Dominant frequency of songs in tropical bird species is higher in sites with high noise pollution European blackbirds in cities preferentially sang higher-frequency song elements, and because frequency and amplitude are tightly linked in their vocal system, these higher-pitched songs also happened to be louder, giving them a double advantage in noisy conditions.24PubMed Central. Bird song and anthropogenic noise: vocal constraints may explain why birds sing higher-frequency songs in cities These vocal adjustments can sometimes compensate for noise, but not always: the effectiveness of the shift depends on the type of noise and even on physiological variation among the receivers trying to hear the signal.25PubMed Central. Noise Source and Individual Physiology Mediate Effectiveness of Bird Songs Adjusted to Anthropogenic Noise

The Emotional Channel

Beyond conveying identity, location, or species membership, vocalizations carry emotional information. In humans, this operates through at least two channels. One is evolutionarily ancient and shared with many other vertebrates: short, nonverbal bursts like screams, laughs, groans, and gasps that are triggered relatively automatically by emotional states and involve limited cognitive processing. These “affective bursts” can be controlled to some degree (you can stifle a laugh or suppress a scream), but doing so requires active effort and a certain level of developmental maturity. The second channel is unique to spoken language: emotional prosody, the way pitch, rhythm, and loudness shift during speech to convey mood or emphasis. Humans use both channels simultaneously. You might say “I’m fine” with prosody that conveys the exact opposite, layering a learned linguistic signal on top of a more primal vocal-emotional one.26PubMed Central. Emotion in Nonverbal Communication: Comparing Animal and Human Vocalizations and Human Text Messages

In other animals, the ancient channel dominates. Farm animal welfare researchers have long recognized that the rate, type, and acoustic features of vocalizations reflect an animal’s internal state, whether it is distressed, content, socially isolated, or in pain.27Applied Animal Behaviour Science. Vocalization of farm animals as a measure of welfare A pig’s squeal during handling or a calf’s repeated calling when separated from its mother are not arbitrary noises; they are readouts of emotional states mediated by the same brainstem-PAG circuits that underlie involuntary vocalizations in humans. Understanding vocalization as an expression of internal state, rather than merely a signal designed for a listener, has practical consequences for animal husbandry, conservation monitoring, and even the development of automated welfare-assessment tools that use acoustic analysis to flag distressed animals in real time.

Leave a Reply

Your email address will not be published. Required fields are marked *