What Is Frequency Analysis?

Frequency analysis is the process of breaking down a complex audio signal into its constituent sine wave components. In the context of speech, this means examining the relative energy present at each frequency across the audible spectrum. The human voice produces a rich harmonic structure, with formants that define vowel sounds and transient bursts that characterize consonants. By performing frequency analysis – typically using a Fast Fourier Transform (FFT) or a spectrogram – engineers and clinicians can visualize exactly where the energy is concentrated and identify anomalies such as excessive resonance, harsh peaks, or muddy buildup.

Understanding the frequency spectrum of speech is essential for audio post‑production, broadcasting, podcasting, and speech‑language pathology. A well‑trained ear can spot problems, but quantitative analysis reveals the precise frequencies that need adjustment. This data‑driven approach separates guesswork from targeted correction.

How to Conduct Frequency Analysis on Speech

Tools and Software

You don’t need an expensive lab. Many free or affordable tools provide real‑time spectral analysis:

  • Praat – a free, cross‑platform program widely used in phonetics and speech therapy. It generates detailed spectrograms and pitch tracks.
  • Audacity – includes a built‑in spectral analysis tool (“Plot Spectrum”) that displays FFT data.
  • iZotope RX – a professional repair suite with a spectrogram viewer, useful for advanced audio cleanup.
  • Overtone Analyzer – emphasizes harmonic content and can highlight frequency buildup.
  • Voxengo SPAN – a free real‑time spectrum analyzer plugin for digital audio workstations.

The Procedure

  1. Capture a clean speech sample in a quiet environment with minimal background noise. Use a cardioid microphone placed consistently (e.g., 6–12 inches from the mouth at a 45° angle).
  2. Open the recording in your analysis tool. Zoom into a representative section – preferably a few seconds of continuous speech containing vowels, consonants, and pauses.
  3. Apply a Fast Fourier Transform or generate a spectrogram. Set the window size (e.g., 2048 samples for a good balance of frequency and time resolution).
  4. Identify peaks or unusual clusters in the frequency domain. Compare against the typical range of speech: the fundamental frequency (F0) for an adult male is ~85–180 Hz, for an adult female ~165–255 Hz. Most intelligibility lies between 300 Hz and 4 kHz.
  5. Look for persistent spikes that deviate from the natural contour. For example, a sharp peak at 800 Hz that remains during both vowels and consonants may indicate a room resonance or a microphone proximity effect.

Interpreting a Spectrogram

A spectrogram plots time on the horizontal axis, frequency (logarithmic scale) on the vertical axis, and intensity as color or brightness. Darker reds or brighter colors indicate higher energy. In speech, you will see horizontal bands (formants) and vertical striations (glottal pulses). Problematic areas often appear as unnatural, stationary bands that do not shift with the voice – these could be hum, buzz, or resonance peaks.

Common Problematic Frequencies in Speech

While every voice is unique, certain frequency ranges consistently cause issues. Recognizing them is the first step toward an effective fix.

Low Frequency Build‑Up (Below 200 Hz)

Excessive energy below 200 Hz results in muddiness, bloated low‑end, and a lack of clarity. This is common when a microphone is too close (proximity effect) or in small rooms with standing waves. Reducing around 100–150 Hz with a gentle high‑pass filter or a narrow cut can clean up the signal without losing natural warmth.

Boominess at 250–400 Hz

This range often contributes to a “boxy” or “honky” tone. It can arise from microphone placement, room reflections, or the speaker’s own vocal tract resonance. A moderate cut (2–3 dB) at around 300 Hz using a parametric EQ can restore neutrality.

Harshness and Nasality (1–3 kHz)

This critical range for speech intelligibility can become harsh when over‑emphasized. Many microphones have a presence peak around 2–4 kHz. If the speaker has a naturally nasal voice, unwanted energy in the 1–2 kHz band may sound piercing. Use a broad cut around 2 kHz, or a dynamic EQ that reduces only when the voice gets loud.

Sibilance and “Ess” Sounds (5–10 kHz)

Sibilance occurs when fricative consonants like “s,” “sh,” “ch,” and “z” produce excessive high‑frequency energy. A de‑esser plugin is the standard tool – it detects sibilant peaks and reduces gain in the 5–8 kHz range. If using a static EQ, a gentle shelf cut above 8 kHz can tame sibilance without dulling the rest of the voice.

Plosive Pops and Thumps (Below 100 Hz)

Plosives (“p,” “t,” “k,” “b”) can create low‑frequency thumps when air hits the microphone diaphragm. These are best fixed at the source with a pop filter and good mic technique. If they still appear, a high‑pass filter set to 80–100 Hz (with a sharp slope) can remove them without affecting vocal fundamental frequencies.

Excessive Brightness or Fizziness (10–15 kHz)

Some voices or microphones produce a metallic or “edgy” quality in the uppermost frequencies. A very gentle high‑frequency shelf cut (starting around 10 kHz) can smooth out the signal. Be cautious – over‑cutting can make speech sound muffled.

Techniques to Fix Problematic Frequencies

Static Equalization (EQ)

Once the problematic frequency has been identified, the simplest fix is to apply a notch or bell filter. Always use a narrow bandwidth (high Q) for problematic resonances to avoid affecting adjacent frequencies. For broader tonal issues, a wider Q is appropriate. The golden rule: cut before you boost. Subtractive EQ preserves headroom and reduces the risk of introducing distortion.

Dynamic Equalization

A static cut or boost applies the same amount of gain reduction regardless of the input level. This can make quiet sections sound thin. Dynamic EQ (or multiband compression) only adjusts the target frequency when its energy crosses a threshold. This is ideal for controlling sibilance, plosives, or harshness that occur only at certain moments. For example, a dynamic EQ set to attenuate 2 kHz by 3 dB when the voice gets louder will reduce harshness without dulling normal speech.

De‑Essers

A de‑esser is essentially a frequency‑selective compressor, typically targeting the 5–8 kHz range. Most modern de‑essers offer controls for threshold, frequency, and reduction amount. For speech, a soft knee and quick attack time (0.5–2 ms) with a release of 20–50 ms works well. If the de‑esser introduces artifacts, try splitting the sibilant band and compressing only the high frequencies.

Noise Gates and Expanders

Background noise can exacerbate frequency problems. A noise gate or a downward expander can clean up the signal between words. Set the threshold so that only low‑level noise (e.g., HVAC hum, distant traffic) is reduced, not the spoken consonants. Place the gate before the EQ in the signal chain so that the gate doesn’t cut off ringing from EQ adjustments.

Multiband Compression

When problematic frequencies are concentrated in a specific band and vary dynamically, a multiband compressor can be more effective than a single de‑esser. For instance, if speech has both a boomy low‑mid (300 Hz) and a harsh high‑mid (2.5 kHz), each band can be compressed independently. Crossovers must be set carefully to avoid “mushy” transitions between bands.

Microphone and Room Adjustments

Often the best fix is not in the EQ but in the capture. Changing microphone type (dynamic vs. condenser), polar pattern (cardioid, figure‑8), or distance can dramatically alter the frequency response. Placing acoustic panels or a portable vocal booth reduces early reflections that cause comb filtering and resonant peaks. Even moving the speaker’s head a few inches can change the balance of frequencies reaching the mic.

Case Study: Diagnosing a Muddly Voice‑Over

Consider a voice‑over recording that sounds unclear and boomy. A spectral analysis reveals a 6 dB peak at 120 Hz and an additional 4 dB peak at 350 Hz. The voice’s fundamental is around 110 Hz, so the 120 Hz peak is likely from microphone proximity effect. The 350 Hz peak adds boxiness. Using a parametric equalizer, a high‑pass filter at 80 Hz (24 dB/octave) reduces the sub‑low rumble, and a narrow cut of 3 dB at 350 Hz (Q = 2) clears the boxiness. After these adjustments, the voice retains its natural fullness but now sounds clear and present. A final A/B comparison confirms the improvement.

For more in‑depth spectral analysis techniques, refer to the Wikipedia article on time–frequency analysis and the Praat software documentation.

Applications Beyond Audio Engineering

Frequency analysis is also a cornerstone in speech‑language pathology. Clinicians use spectrograms to assess articulation disorders, voice disorders, and resonance problems. For example, a child with a lateral lisp may show a broad noise band between 5–10 kHz on /s/ sounds. By comparing the spectrogram against a typical /s/ (which has a sharp, narrow band near 8 kHz), the therapist can monitor progress. Vocal pathologies such as vocal fold nodules often manifest as increased noise and irregular formant structure in the 1–4 kHz range.

In audio forensics, frequency analysis helps authenticate recordings or isolate specific voices from background noise. The same principles apply to music production: removing guitar amp hiss, controlling vocal resonant peaks, and equalizing instrument bleed all rely on understanding the frequency spectrum.

Pitfalls to Avoid

  • Over‑EQing: Applying too many cuts can leave the speech sounding hollow and unnatural. Aim for no more than three to five EQ moves per voice.
  • Ignoring the time domain: Frequency analysis alone doesn’t reveal transient issues like clicks or pops. Always listen in addition to looking at the spectrum.
  • Boosting without reason: Adding gain to a frequency that is not deficient will only increase noise and distortion. Boost only when you have a clear goal (e.g., adding airiness with a 12 kHz shelf).
  • Using steep filters unnecessarily: While high‑pass filters are useful, a 48 dB/octave slope can create phase shifts that affect transient clarity. Start with 12 dB/octave if the problem is minor.

Conclusion

Frequency analysis transforms subjective criticism – “the voice sounds muddy” – into objective data – “there is a 5 dB peak at 150 Hz and a 3 dB dip at 3 kHz.” Armed with this information, you can apply surgical EQ, dynamic processing, or acoustic treatment to achieve clear, natural, and intelligible speech. Whether you are a podcast producer, an audio engineer, or a speech therapist, mastering the art of spectral reading and correction will significantly elevate the quality of your work.

For further reading on equalization techniques, see Sound On Sound’s guide to EQ fundamentals and the American Speech‑Language‑Hearing Association’s voice disorders portal. With consistent practice, frequency analysis will become an intuitive part of your workflow.