VoiceLab Article

Human Voice Frequency Map: The definitive guide to mixing voices

A full psychoacoustic map of the human voice, from subsonic weight to ultrasonic air, designed for producers who mix with intention.

Executive summary

Premium VoiceLab encyclopedia for producers on vocal frequency bands, emotional function, EQ, saturation, compression, plugin choices and ear training.

How this research informs production

This publication is written for producers, creative directors, AI voice teams and brand leads who need to evaluate a voice before committing to a campaign, dataset or long-form narration workflow.

The practical value is not academic distance. It is a clearer production conversation: what the voice can carry, how much intimacy or authority it offers, where mastering should stay transparent, and which recording conditions protect the signature voice.

For campaign, documentary, corporate, luxury and AI voice work, this framing turns acoustic observations into usable decisions about casting, briefing, recording, approvals and final delivery.

Table of contents

  1. How to use this frequency map
  2. 20-80 Hz
  3. 80-250 Hz
  4. 250-800 Hz
  5. 800 Hz-2.5 kHz
  6. 2.5-6 kHz
  7. 6-12 kHz
  8. 12-20 kHz
  9. Examples by production format
  10. Text diagrams for the frequency map
  11. Plugin recommendations by task
  12. How to hear frequencies and train the ear
  13. Premium vocal mixing workflow

How to use this frequency map

A voice frequency map is not a recipe book. It is a listening architecture. The same 3 dB move can make one voice intimate, another aggressive and another completely artificial. The producer must read the band, the source, the microphone, the room, the performance and the delivery format at the same time. Premium vocal mixing begins when frequency becomes psychology.

This guide separates the human voice into seven working zones: 20-80 Hz, 80-250 Hz, 250-800 Hz, 800 Hz-2.5 kHz, 2.5-6 kHz, 6-12 kHz and 12-20 kHz. For every band, the important question is not only what it sounds like, but what the listener believes because of it. Does the voice feel trustworthy, too close, expensive, nervous, present, tired, synthetic or alive? That is the real map.

Use the ranges as overlapping territories rather than hard borders. A bass voice, a bright female lead, a Spanish commercial read, a Catalan audiobook, an English cinematic narration and a TTS voice do not place emotion in exactly the same point. Still, the psychoacoustic logic repeats: low frequencies give ground, low mids give flesh, mids give language, presence gives contact, air gives finish and the top octave gives halo.

20-80 Hz

This band is not where speech meaning lives, but it changes the emotional floor of a recording. In professional voice it is mostly felt as room movement, mechanical vibration, mic stand noise, plosives, transport rumble and low stage energy. A little controlled depth can make a cinematic narrator feel larger than the speakers. Too much makes the voice heavy before it becomes intimate. Technically, most voiceover, podcast, audiobook and TTS work should be filtered here because the band consumes headroom without improving intelligibility. Use a high-pass filter with intention: often 60-80 Hz for spoken voice, lower for a very deep cinematic voice, higher for thin mobile playback. Excess creates boom, limiter pumping, false proximity and muddy reverbs. Defect rarely matters, except in trailer or cinematic narration where removing everything can make the voice feel detached from the room.

Typical settings: 12 or 18 dB per octave high-pass around 55-80 Hz, with automation if plosives vary. Saturation should be minimal; if used, choose transformer style color that creates harmonics above the fundamental rather than adding more sub. Compression should avoid reacting to this range unless the goal is special effect. Sidechain high-pass the detector in broadband compressors so sub energy does not pull down the whole voice.

80-250 Hz

The 80-250 Hz range gives the voice physical confidence. It is chest, mass, gravity, warmth and the feeling that a person has a body behind the words. In commercial voiceover it can create authority without shouting. In podcast it gives proximity and trust. In audiobooks it lets the narrator sit with the listener for hours. In cinematic narration it can make a line feel inevitable. Technically, this area carries low fundamentals, proximity effect and part of the room signature. Excess creates boom, pillow tone, mouth too close to the microphone, slow transients and a mix that becomes large but not clear. Defect makes the voice thin, nervous, papery or disconnected from the image.

Typical settings: small shelves or bell moves around 100-180 Hz, often plus dynamic control when proximity changes between phrases. Saturation can be tube, tape or transformer style, but lightly, because harmonic enrichment in this zone can become mud quickly. Compression works well with slow attack and medium release when you want to preserve chest movement; for podcast consistency, use gentle leveling before stronger final control. Avoid crushing this range with fast compression because the voice loses breath and scale.

250-800 Hz

This is the most misunderstood voice band. Emotionally it contains warmth, wood, humanity and the sense that the voice is made of a real vocal tract. Technically it also contains boxiness, room buildup, booth tone, nasal lower resonance and the veil that makes a loud voice feel covered. Many producers ask for more presence when the real problem is not 3 kHz, but too much 350-600 Hz. Excess in this zone creates cardboard, congestion, muffled diction, low-mid fatigue and a voice that feels close but not intimate. Defect creates a scooped, expensive-in-solo but poor-in-context voice that lacks continuity.

Typical settings: subtractive EQ around 300-500 Hz for mud, narrower cuts near 550-750 Hz for box or tube resonance, and dynamic EQ if the buildup appears only on certain vowels. Saturation should be chosen carefully: tape can smooth this area, but dense tube saturation may thicken it too much. Compression should be transparent and frequency-aware. Multiband compression can restrain low-mid bloom without flattening the whole performance. In audiobook and podcast work, this band decides whether the listener stays relaxed or becomes tired.

800 Hz-2.5 kHz

This range is the architecture of language. Emotionally it gives direction, intention and the sense that the speaker is choosing words, not merely producing sound. Technically it carries vowel definition, forward tone, much of consonant support and the shape that lets a phrase read on small speakers. Excess creates honk, nasal push, telephone hardness and a voice that seems to lecture the listener. Defect makes the voice polite but distant, glossy but vague, or buried under guitars, synths and ambience.

Typical settings: broad presence preparation between 1 and 2 kHz, cautious cuts around nasal nodes, and automation when a phrase needs to step forward. Saturation can be excellent here if it is harmonic and not brittle: console, tape or subtle triode color can make words easier to read without simple EQ boost. Compression should preserve microphrasing. Too much fast compression in this band makes every syllable equally urgent. Dynamic EQ or multiband control is often more elegant than one static boost.

2.5-6 kHz

This is the presence corridor. It contains consonant attack, articulation, lip edge, bite, the first feeling of proximity and the zone where a voice can cross a dense arrangement without more fader. Around 3.3 kHz often sits the Presence Anchor, the point where the voice starts to appear emotionally in front of the listener. Excess creates aggression, cheap advertising pressure, ear fatigue, pain on headphones and a sense that the speaker is pushing. Defect makes the voice dull, slow, behind the music and emotionally undecided.

Typical settings: small broad lifts around 3-4 kHz only after cleaning 250-800 Hz; dynamic cuts around 3.5-5 kHz when certain words stab; careful de-essing if consonants and sibilance overlap. Saturation should be restrained and high quality, because hard clipping in this range quickly becomes sandpaper. Compression should be phrase-aware. Use dynamic EQ, split-band de-essing or multiband compression that moves only when the voice becomes sharp. For TTS and repeated interface voices, this zone must be clear but never punitive.

6-12 kHz

The 6-12 kHz band gives air, polish, breath, saliva, intimacy, modern brightness and perceived resolution. Emotionally it can make a voice feel premium, clean and expensive, but it can also reveal anxiety, dryness, editing artifacts and synthetic texture. Technically it holds sibilance, lip noise, de-esser behavior, top-end compression artifacts and the final sheen that makes a vocal sit in a finished production. Excess creates sibilance, glass, fatigue, brittle consonants and a voice that looks bright but feels less human. Defect creates a covered, old, blanket-like or over-de-essed voice.

Typical settings: de-ess before adding air, lift with a gentle shelf only if the mouth and recording support it, and check headphones, phone speaker and quiet playback. Saturation can be useful if it creates soft upper harmonics rather than fizz; tape and refined exciter tools can work, aggressive enhancers can damage the illusion. Compression should be extremely careful. Fast broadband compression often exaggerates sibilance. Split-band control around 6-9 kHz and separate air shaping above that are usually cleaner.

12-20 kHz

Most speech information is below this band, but 12-20 kHz affects perceived openness. Emotionally it is halo, space, clean glass and the suggestion that the recording has air around it. Technically it carries the very top of breath, converter tone, noise shaping, oversampling artifacts, hiss and the polish created by mastering chains. Excess is rarely useful in long-form voice. It can create hiss, false luxury, brittle streaming encodes and listener fatigue even when the meter looks harmless. Defect is often acceptable for podcast or audiobook, but in premium commercial and pop vocal work a complete absence can make the voice feel small or dated.

Typical settings: very gentle shelves above 12 or 14 kHz, only after sibilance has been solved. Saturation should be oversampled and elegant; otherwise aliasing can turn the halo into grit. Compression is usually unnecessary here except for controlling hiss or exaggerated air with dynamic EQ. Always compare after codec conversion, because some bright masters collapse when encoded.

Examples by production format

Commercial voiceover usually needs a clean low end, controlled low mids and a deliberate Presence Anchor. A car campaign can use 100-180 Hz for confidence and 3-4 kHz for command, while a luxury fragrance spot may reduce low-mid density, keep the 2.5-6 kHz zone silky and add only a controlled breath shelf. The mistake is making every advert loud, bright and compressed. Premium advertising sounds expensive because it has restraint.

Podcast voice asks for comfort before spectacle. Filter 20-80 Hz, control 120-220 Hz proximity, remove box around 300-600 Hz and avoid a static 4 kHz spike that becomes exhausting after twenty minutes. Audiobook mixing is even more conservative: the listener needs hour-long continuity, so harsh presence and artificial 12 kHz air are more dangerous than a slightly modest top end. Narration must disappear as technology and remain as person.

TTS and voice systems require consistency without punishment. The 2.5-6 kHz band must carry intelligibility, but repeated prompts turn small sharpness into real fatigue. Use dynamic EQ on the presence band, soft de-essing, controlled 6-12 kHz and almost no dramatic low-end enhancement. Pop vocals are different: they can accept more saturation, parallel compression and air because they compete with drums, bass and synths. Even there, the vocal must keep vowel truth and not become only a bright object.

Cinematic narration uses the whole map. It may keep more 80-250 Hz body, less aggressive 3-5 kHz, deep silence between phrases and a carefully designed 12-20 kHz halo. The voice should feel larger than ordinary speech but not detached from the human mouth. When the narration supports picture, the best EQ decision is often the one that makes the scene believable rather than the voice impressive in solo.

Text diagrams for the frequency map

Diagram 1: imagine the voice as a building. 20-80 Hz is the ground vibration under the building. 80-250 Hz is the foundation slab. 250-800 Hz is the wooden structure. 800 Hz-2.5 kHz is the corridor that lets language move. 2.5-6 kHz is the front door where the listener meets the speaker. 6-12 kHz is the glass and light. 12-20 kHz is the sky above the roof. If the foundation is too large, the door disappears. If the glass is too bright, nobody trusts the room.

Diagram 2: imagine a horizontal line from left to right: weight, body, word, presence, air, halo. A good vocal mix is not a smile curve; it is a balanced sentence. The listener should travel from body to meaning without noticing the technical path. When you boost a band, ask which word in that diagram becomes louder emotionally.

Diagram 3: imagine three vertical layers. The bottom layer is stability: 80-250 Hz plus controlled 250-500 Hz. The middle layer is identity: 500 Hz to 2.5 kHz. The top layer is contact: 2.5 kHz upward. If the top layer is mixed before the bottom is cleaned, the voice becomes bright dirt. If the bottom is removed before identity is understood, the voice becomes expensive but empty.

Plugin recommendations by task

For surgical EQ, FabFilter Pro-Q 3, Kirchhoff-EQ, DMG Equilibrium and TDR Nova are strong choices. Use them for high-pass filtering, low-mid cleanup, nasal nodes and dynamic presence control. For resonance management, oeksound soothe2, Waves Silk Vocal or smart dynamic EQ can help, but the goal is not to erase personality. Remove what blocks the phrase, not what makes the voice human.

For saturation, Soundtoys Decapitator, FabFilter Saturn 2, UAD Studer, Softube Tape, Black Box HG-2 and transformer-style console emulations can add density. Use them differently by band: harmonic support in 80-250 Hz, careful warmth in 250-800 Hz, articulation density in 800 Hz-2.5 kHz and almost no aggressive distortion above 6 kHz. Oversampling matters when the top end is exposed.

For compression, classic 1176-style control can bring urgency, LA-2A-style leveling can smooth narration, R-Vox-style tools can solve quick broadcast consistency, and Pro-MB or C6-style multiband compression can restrain specific regions. For de-essing, FabFilter Pro-DS, Oxford SuprEsser, Weiss Deess and RX tools are reliable when set by listening, not by meters. The best chain is often simple: clean EQ, light saturation, phrase compression, dynamic presence control, de-ess, final tone.

How to hear frequencies and train the ear

Ear training for voice begins with subtraction. Take a clean spoken recording and create seven EQ bands. Boost one band by 6 dB for ten seconds, name the emotional change, then return to zero and cut the same band by 6 dB. Do not start by memorizing numbers. Start by naming perception: ground, chest, mud, word, bite, air, halo. When the word is clear, the number becomes useful.

Practice at three monitoring levels. At low volume, presence and intelligibility reveal themselves honestly. At medium volume, body and low mids become easier to judge. At high volume, harshness and sibilance become obvious, but do not make final decisions there because the ear protects itself. Repeat the same exercise on monitors, closed headphones, earbuds and phone speaker. The band that survives every playback system is usually the real problem.

Use pink noise and references only as calibration, not as taste. Compare a podcast voice you trust, an audiobook you can listen to for an hour, a pop vocal that cuts without pain, a luxury commercial read and a cinematic narration. Sweep slowly, but do not mix while sweeping. Sweeping exaggerates problems. Real decisions come from bypassed comparison, level matching and asking one question: does the voice become more believable?

A useful weekly drill is the seven-band diary. For each mix, write one sentence per band before touching plugins. 20-80 Hz: what noise is present? 80-250 Hz: does the body support or inflate? 250-800 Hz: is there warmth or veil? 800 Hz-2.5 kHz: do words read? 2.5-6 kHz: is the speaker present or sharp? 6-12 kHz: is the air human or cosmetic? 12-20 kHz: is the halo useful or only hiss? This habit turns EQ from reaction into listening discipline.

Premium vocal mixing workflow

A premium vocal workflow starts before inserting plugins. First decide what the voice must do inside the production: sell, accompany, teach, confess, seduce, guide, narrate, sing or become a system voice. That verb changes the frequency map. A commercial voice may need faster contact in the 2.5-6 kHz zone. A podcast may need relaxed low mids and less top-end theatre. An audiobook may need continuity above all. A pop vocal may need saturation and parallel control. A cinematic narration may need body, silence and scale. The same chain cannot serve every intention.

Step one is technical hygiene. Remove 20-80 Hz noise, edit plosives, correct mouth clicks only when they distract, and gain-stage phrases so the compressor does not become a rescue tool. Step two is identity EQ. Listen to 80-250 Hz and 250-800 Hz before touching presence. If the body is unstable, the voice will never feel expensive. If low mids are cloudy, any bright boost will sit on top of mud. Step three is language EQ: 800 Hz-2.5 kHz must let words read without turning the speaker into a megaphone.

Step four is presence design. Search for the Presence Anchor with the whole mix playing. Do not solo the voice for too long. A boost that sounds impressive in solo can become painful against cymbals, guitars, synths or room ambience. Use automation before static EQ when the problem is phrase-dependent. Use dynamic EQ when a band is right most of the time but wrong on certain vowels or consonants. The most elegant vocal mix is rarely the one with the biggest processing chain; it is the one where every processor has a perceptual reason.

Step five is density and control. Saturation should answer a question: does the voice need harmonic body, articulation density, softened peaks or emotional grain? Compression should answer another: does the performance need leveling, urgency, intimacy, stability or energy? If you cannot name the reason, bypass the plugin. For spoken voice, serial compression often beats one aggressive compressor: a gentle leveler, then a faster peak controller, then final bus control. For sung pop vocals, parallel compression and saturation can create size while the dry vocal keeps expression.

Before printing the mix, make a frequency pass with the arrangement muted only for short checks, then immediately return to context. Mark one decision per zone: filter, support, clean, define, present, polish or leave untouched. This prevents overmixing. A professional chain does not need to process every band. Sometimes the most senior decision is to protect a band because it already carries the right emotion. Silence, bypass and level matching are part of the map.

The final step is translation. Check the voice at low volume, on headphones, on a phone, through the full mix, after limiting and after codec conversion. If the vocal only works loud, it is not mixed; it is inflated. If it remains believable when quiet, clean when encoded and emotionally clear inside the arrangement, the frequency map is doing its job. A premium mix is not the absence of problems. It is the presence of a convincing human center. The listener should feel the voice clearly before noticing the processing.

Key findings

Technical conclusions

The definitive voice frequency map is not a list of boosts and cuts. It is a way of connecting frequency to human belief. A premium vocal mix gives the listener body, language, presence, air and truth in the right proportion for the format.

Production questions answered

How should a producer use this Voice Lab article?

Use it as a voice-direction reference before casting or recording. It clarifies acoustic identity, mastering choices, AI voice relevance and the kind of brief JOIA needs to deliver useful takes.

Can the findings support AI voice or dataset planning?

Yes. The findings help define consistency, vocal identity, prompt design, consent-aware usage and review criteria before a TTS, voice cloning or conversational AI recording session.

What is the commercial value of the research?

It gives agencies, brands and production teams a shared language for tone, warmth, clarity, authority, intimacy and broadcast finish, which reduces vague feedback during recording.

Apply this research to a voice project

Send a script, campaign context or AI voice requirement and ask for voice direction, recording availability or usage guidance. For a quote, availability check or directed session, include script length, usage, market and deadline.

Discuss a voice project