When recording speech—whether for a podcast, voice-over, broadcast, or video narration—one of the most common challenges engineers face is harshness in the natural timbre. Those piercing “s” sounds, brittle high-mid transients, or aggressive consonants can distract listeners and make a recording sound unpolished. While equalization can help, it often removes too much of the voice’s character. Multi-band compression offers a more intelligent solution: a dynamic, frequency-aware way to tame problematic areas without sacrificing clarity or tone. By splitting the signal into separate bands and compressing each independently, you can keep the speech full, present, and smooth—even when the performance varies wildly in energy.

What Multi-Band Compression Really Does

At its core, multi-band compression works by first sending the audio signal through a series of crossover filters that divide the spectrum into two, three, four, or more bands. Each band is then fed into its own compressor, with independent controls for threshold, ratio, attack, release, and even make-up gain. The processed bands are recombined at the output to produce the final signal.

This contrasts sharply with a conventional single-band compressor, which reacts to the whole mix. A sharp sibilant peak in the upper midrange, for example, can cause the compressor to pull down the entire signal—including desirable low-end fullness and mids—making the voice sound thin or distant. Multi-band compression avoids this by confining the reduction to the specific band where the problem occurs. The result is surgical dynamic control that preserves natural voice characteristics while still taming harshness.

Typical Band Splits for Speech

Most multi-band compressors allow you to set the crossover frequencies. For speech, a common starting point is:

  • Low band: 20–300 Hz (rumble, proximity effect, fundamental frequencies)
  • Low-mid band: 300 Hz–2 kHz (body, nasal fullness, vowel shape)
  • High-mid band: 2–6 kHz (presence, edge, sibilance consonants)
  • High band: 6–20 kHz (air, sibilance fricatives, harsh transients)

For taming harshness, focus primarily on the high-mid band around 4–8 kHz, and sometimes the high band above 8 kHz. The precise location of harshness depends on the speaker’s voice, microphone, and recording environment.

Why Speech Develops Harsh Frequencies

Harshness in speech usually originates from two sources: the sound source itself (e.g., fricatives like “s,” “sh,” “z,” and “ch”), and the recording chain. A bright condenser microphone placed too close to the mouth can accentuate sibilant energy above 5 kHz. A preamp or converter with harsh high-frequency distortion can also add grit. Even a well-recorded voice can sound brittle in untreated rooms due to comb filtering and reflections.

Female voices often have more energy in the 4–7 kHz region than male voices, making them more prone to sibilance issues. Conversely, loud male voiceovers can produce aggressive “punch” in the 2–5 kHz range that sounds fatiguing. Multi-band compression allows you to handle each speaker’s unique frequency fingerprint without a blanket EQ cut that would remove their natural presence.

Key Benefits of Multi-Band Compression for Speech

While single-band compression and de-essing are valuable tools, multi-band compression provides several advantages that make it indispensable for professional speech processing:

  • Frequency-specific dynamic control: You can reduce only the spikes that cause harshness, leaving the fundamental character of the voice untouched.
  • Preservation of low-end weight: Taming a high-mid harsh transient won’t thin out the chesty low frequencies or compromise the warmth of the voice.
  • Consistent timbre across varying dynamics: A speaker who moves closer and farther from the microphone will cause changes in proximity effect and sibilance levels; multi-band compression adjusts per frequency range to keep the tone stable.
  • Reduced listener fatigue: Harsh frequencies quickly tire the ear. Controlling them smoothly allows listeners to focus on the content rather than the recording quality.
  • Less need for aggressive EQ: With precise dynamic control, you can often avoid heavy high-frequency cuts that would dull the voice. The result is a more open and natural sound.

How to Use Multi-Band Compression Effectively

Applying multi-band compression to speech requires a methodical approach. Rushing into settings can make the voice sound phasey, dull, or artificial. Follow these steps to get clean, transparent results.

1. Identify the Problem Frequencies

Use a spectrum analyzer (many multi-band compressors include one) to pinpoint where harshness lives. Playback the loudest sibilant moments and look for peaks or energy clusters between 4 kHz and 8 kHz. Sometimes harshness extends upward to 10–12 kHz with sharp fricatives like “s” and “f.” Also listen for any “honky” or “tinny” resonances around 1–3 kHz that might need attention.

2. Set Crossover Frequencies

Based on your analysis, define the band edges. If harshness is mostly around 5 kHz, split your bands so that one band covers roughly 3 kHz to 7 kHz. Avoid placing a crossover directly on the peak—this can cause phase interference. Instead, keep the crossover frequencies in less active areas (e.g., 2 kHz, 700 Hz, 9 kHz). Most engineers use three or four bands for speech.

3. Adjust Threshold and Ratio

Start by soloing the band you intend to compress. Lower the threshold until you see 3–6 dB of gain reduction on the worst moments. Use a low-to-moderate ratio (2:1 to 4:1). Higher ratios can sound unnatural; the goal is to smooth out peaks, not squash them. For extremely spiky sibilance, a ratio of 5:1 can work briefly, but be prepared to adjust attack and release to avoid pumping.

4. Set Attack and Release Times

For speech, fast attack times (1–5 ms) are usually needed to catch harsh transients. If the attack is too slow, the peak will pass through before compression engages. Release times should be quick enough to let the compressor recover between syllables but not so fast that it causes distortion. A release of 50–80 ms is a good starting point; adjust based on the speed of the performer’s speech. For very fast talkers, shorter releases (30–50 ms) may be necessary.

5. Adjust Makeup Gain and Listen in Context

After setting compression, use the band’s makeup gain to bring the level back to where it was before processing (or slightly below, to ensure a natural balance). Then listen to the full mix of all bands. Compare bypass vs. engaged. The compressed version should sound smoother, less piercing, but not dull or distant. If the voice feels “covered” or the high end seems missing, you are compressing too aggressively—back off the ratio or raise the threshold.

Advanced Techniques

Once you are comfortable with basic multi-band compression, several more advanced tactics can elevate your speech processing.

Parallel Multi-band Compression

Mix the compressed band with a dry version of itself. This preserves the natural attack and air while still reducing harsh peaks. Many multi-band plug-ins allow a “mix” control per band. A blend of 50–70% wet can yield a very transparent effect, especially for subtle sibilance control.

Using Side-Chain Inputs

Some multi-band compressors let you trigger the compression in one band using a different frequency range. For speech, you could side-chain the high band from the low band to avoid compressing sibilance during loud low-frequency vowels. This keeps the voice full without over-processing.

Dynamic EQ vs. Multi-band Compression

Dynamic EQ is a close cousin that reduces a specific frequency when it exceeds a threshold, using a parametric bell curve. Multi-band compression uses a full-band compressor per band, which can sometimes cause broader frequency changes. For wide harshness around 5 kHz, multi-band compression is often more effective; for very narrow, resonant peaks, dynamic EQ may be cleaner. Many engineers use both: dynamic EQ to notch a single resonance, and multi-band compression for general harshness control.

Practical Tips for Better Results

  • Always bypass one band at a time to hear what each is contributing. It’s easy to over-process multiple bands without realizing one band is doing all the work.
  • Monitor on multiple playback systems: What sounds smooth on studio monitors may be harsh on laptop speakers or earbuds. Check your mix on a phone speaker, headphones, and a car stereo.
  • Use a pre-compression high-pass filter: Remove rumble below 80 Hz before the multi-band compressor so that the low band doesn’t react to microphone handling noise or HVAC rumble.
  • Combine with a dedicated de-esser for heavy sibilance: If sibilance is extreme, apply a gentle de-esser (or a dynamic EQ with a bell around 6–8 kHz) before the multi-band compressor to take the edge off, then use multi-band compression for overall smoothing.
  • Adjust bandwidths carefully: Using too wide a band (e.g., 2–12 kHz) will affect too much of the voice. Narrower bands (e.g., 4–6 kHz) give more precise control but can cause phase issues if the crossover slopes are too steep. Good plug-ins use Linkwitz-Riley filters that sum flat, but always check.
  • Purposely over-compress while setting up, then dial back: This technique helps you hear exactly what the band is doing. Once you identify the effect, reduce the ratio or raise the threshold until it sounds natural.

Common Pitfalls to Avoid

Even experienced engineers can fall into traps with multi-band compression. The most frequent issues include:

  • Over-compression leading to a “glassy” or “plastic” sound: Too much gain reduction across multiple bands makes the voice sound processed and unnatural. Always compare against the original.
  • Phase cancellation at crossover frequencies: If the plug-in uses linear-phase filters, you might notice preringing or smearing on transients. Minimal-phase filters can cause group delay differences. Listen for any comb-filtering on the word “s” or “t.” If you hear it, adjust crossover points slightly or try a different mode.
  • Pumping or breathing: When the release time is too short, the compressor recovers so quickly that it “pumps” in time with the speech rhythm. Increase the release or reduce the ratio.
  • Ignoring the low bands: Low-frequency fluctuations from plosives or proximity effect can cause the low band to compress, muddying the voice. Use a high-pass filter on the sidechain of the low band to prevent it from reacting to breath noise.

Workflow Integration in Post-Production

Multi-band compression doesn’t replace other processing; it enhances it. A typical speech chain might look like:

  1. High-pass filter (remove rumble)
  2. Subtractive EQ (remove mid-range resonances and low-mid muddiness)
  3. Multi-band compressor (tame harsh frequencies and smooth dynamics)
  4. Gentle broadband compressor (even out overall level)
  5. Limiter (catch any remaining peaks)
  6. Warmth or saturation (optional, to restore perceived presence)

When placed after subtraction EQ, the multi-band compressor only needs to deal with real harshness—not resonances that were already reduced. A gentle broadband compressor after multi-band compression ensures consistent output without the multi-band unit having to do all the dynamic work.

Conclusion

Multi-band compression is an essential tool for any audio engineer working with speech. By isolating harsh frequency regions and compressing them independently, you can achieve a level of control that equalization and single-band compression alone cannot match. The result is speech that retains its natural expressiveness, presence, and clarity without listener fatigue. Whether you are producing a podcast, recording a voice-over for a commercial, or cleaning up dialogue for video, learning to wield multi-band compression effectively will dramatically improve your final product. Practice with different voices, microphone setups, and material to build an intuitive sense of how each band interacts with the acoustic environment. With time, you’ll be able to hear a harsh frequency and know exactly where to set your crossover, threshold, and ratio—delivering professional, polished speech every time.

For further reading, explore this Sound On Sound guide on multi-band compression, or check out FabFilter’s in-depth tutorial on set-up and creative use. For a deeper dive into speech-specific processing, see iZotope’s speech compression tips.