audio-production-techniques
Methods for Managing Plosive and Sibilant Sounds in Dialogue Tracks
Table of Contents
Methods for Managing Plosive and Sibilant Sounds in Dialogue Tracks
Dialogue is the backbone of most audio productions, from podcasts and audiobooks to film and television. Achieving clear, natural-sounding speech requires careful attention to two common but troublesome artifacts: plosive and sibilant sounds. Plosives produce low-frequency pops, while sibilance creates harsh, hissing noises. Both can break immersion and fatigue listeners. This guide covers proven techniques for preventing and reducing these issues during recording and post-production, helping you deliver professional dialogue tracks.
Understanding Plosives and Sibilance in Dialogue
Plosive sounds occur when air is suddenly released from the mouth, hitting the microphone diaphragm with a burst of energy. The most problematic plosives are P, B, T, D, and K. Sibilance, on the other hand, consists of high-frequency fricatives like S, Sh, Z, Ch, and J. These sounds can become exaggerated due to microphone proximity, room acoustics, or vocal characteristics. Understanding the physical and acoustic causes is the first step toward effective management.
Preventive Recording Techniques
Addressing plosives and sibilance at the source saves hours of post-production work. Proper microphone technique and environment setup are critical.
Microphone Placement and Polar Patterns
Position the microphone slightly above or below the speaker’s mouth (off-axis) rather than directly in front. A 45-degree offset reduces the direct air blast from plosives while still capturing clear speech. The microphone should be 6–12 inches away from the speaker. Closer placement boosts proximity effect, which can exaggerate low-frequency plosives and also increase sibilance due to higher overall level. Using a cardioid or hypercardioid polar pattern can help reject off-axis noise, but be aware that hypercardioid patterns have a rear lobe that may pick up reflections. For sibilance reduction, dynamic microphones like the Shure SM7B or Electro-Voice RE20 are often preferred because they have a naturally smoother high-frequency response compared to large-diaphragm condensers. However, many voice actors use condensers with added control.
Pop Filters and Windscreens
A pop filter is a simple, inexpensive tool that diffuses the air burst from plosives. Dual-layer mesh filters are most effective. For field recording, foam windscreens reduce wind noise and can soften plosives, though they are less effective than pop filters. Always position the pop filter a few inches from the microphone; the filter itself should be 2–4 inches from the mic capsule for best results.
Speaker Technique and Monitoring
Encourage talent to speak slightly to the side of the mic or to back off during plosive consonants. Some voice actors learn to “de-plode” by relaxing the lips and reducing puff on P and B sounds. Monitoring with headphones allows immediate detection of problems—if you hear a pop or hiss, adjust immediately.
Post-Production Techniques for Plosive Control
When plosives survive the recording stage, several editing tools can minimize them without destroying the natural sound of the dialogue.
High-Pass Filtering and EQ
Plosives generate energy below 100–150 Hz. Applying a high-pass filter (HPF) around 80–120 Hz can reduce the lowest thumps while preserving vocal body. A steep roll-off (24 dB/octave) works best for removal. Conversely, for a few lingering plosives, a narrow EQ cut around 50–100 Hz can be automated. Avoid cutting too much low end, or the voice will sound thin.
Clip Gain and Fades
A manual but precise method: locate the plosive waveform—it often appears as a large, asymmetrical spike with a low-frequency tail. Reduce the clip gain by 3–6 dB for that syllable using volume automation or clip gain, then apply a short fade-in (1–5 ms) at the start of the syllable to soften the attack. Crossfading with neighboring syllables prevents clicks. This technique preserves the breath and speech dynamics while eliminating the pop.
Spectral Editing
Advanced spectral editors like iZotope RX allow you to see plosives as bright, low-frequency blobs. Using the “Spectral Repair” module, you can replace the plosive region with surrounding noise or attenuate it. The “De-plosive” module in RX works by detecting plosive events and applying automated reduction. Many DAWs also offer third-party plugins with similar functionality.
Post-Production Techniques for Sibilance Reduction
Managing sibilance requires careful balancing: over-processing leads to lisping or dull speech, while under-processing leaves harshness. The goal is to tame the high-frequency energy without affecting the natural sibilant consonants.
De-essing Basics
De-essers are compressors that act only on a user-defined frequency band, usually 5–10 kHz. The key parameters are threshold, ratio, and frequency range. Set the frequency by sweeping a narrow EQ boost until the sibilance becomes obvious, then reduce gain at that freq with the de-esser. A ratio of 2:1 to 4:1 works for most material. Listen for artifacts like “lisping” (overly dull S’s) or “shushing” (too much reduction on Sh). Modern de-essers like the FabFilter Pro-DS and Waves Sibilance allow split-band processing for transparent results.
Multiband Compression
For more control, use a multiband compressor on the high-frequency band (5–15 kHz). This applies dynamic reduction only when sibilance exceeds a threshold. Adjust attack and release times: faster attack (1–5 ms) catches the sibilant burst without affecting the rest of the speech. A softer knee can smooth the transition. Multiband compression is especially useful for voice actors with naturally sibilant speech because it adapts to varying intensity.
Manual Spectrum Editing
For stubborn sibilance, copy the problematic S or Sh sound to a new track, apply heavy high-frequency filtering, and crossfade the processed track with the original. You can also use spectral editing tools to isolate and reduce sibilant bands in the frequency domain. In iZotope RX, the “Spectral De-ess” module uses machine learning to target only sibilant phonemes, leaving other sounds untouched. This method yields natural results but is more time-consuming.
Advanced Workflow Considerations
Combined Handling of Plosives and Sibilance
Dialogue tracks often contain both issues. A workflow may be: start with a high-pass filter to reduce low-end plosive rumble, then apply de-essing while monitoring in context. After level adjustments, manually edit any remaining plosive pops using clip gain and fades. Finally, listen on multiple playback systems (headphones, nearfield monitors) to catch hidden artifacts. Always process in the context of the full mix, as music and background noise can mask or reveal problems.
Room Acoustics and Treatment
A well-treated room reduces the need for aggressive post-processing. Reflections from hard surfaces can amplify both plosive pops and sibilant hisses. Use absorption panels to dampen high frequencies and bass traps for low-frequency plosives. Portable isolation shields help when recording in untreated spaces. Keep the recording area quiet: air conditioning or other noise can mask sibilance until you try to clean it.
Conclusion
Managing plosive and sibilant sounds is an essential skill for any audio professional working with dialogue. The most effective approach combines good recording techniques—proper microphone placement, pop filters, and speaker awareness—with versatile post-production tools like high-pass filters, manual clip gain, de-essers, multiband compressors, and spectral editors. With practice, you can reduce these distractions while retaining the natural quality and character of the voice. Ultimately, clean, intelligible dialogue keeps listeners engaged and upholds the production value of any project. For further reading, check out guides on microphone placement for voice and advanced de-essing techniques from iZotope.