Understanding the Most Frequent Audio Flaws in Podcasts

Before you can fix a problem, you need to recognize it. Podcast audio issues typically fall into a handful of categories, each with its own causes and solutions. By learning to identify these flaws in your recordings, you'll be able to apply the right correction during mixing and avoid wasting time on guesswork.

  • Uneven volume levels – drastic loudness shifts between speakers, segments, or even within a single sentence. This forces listeners to constantly adjust their playback volume.
  • Background noise – persistent hums, hisses, room echo, or random clicks and bumps. These distract from the content and make speech harder to follow.
  • Clipping and distortion – harsh, crackling artifacts that occur when audio levels exceed 0 dBFS (full scale). Digital clipping is often irreversible.
  • Inconsistent tonal quality – one voice sounds muddy while another is thin and harsh, often due to different microphones or recording environments.
  • Sibilance and plosives – harsh “s” and “sh” sounds (sibilance) and explosive “p,” “b,” and “t” sounds (plosives) that can be fatiguing or even painful to listeners.

These problems aren't just technical annoyances. Poor audio quality drives listener drop-off. Research shows that even slightly distracting sound can cause audience members to stop listening within seconds. By systematically addressing these issues in the mix, you build trust with your audience and make your podcast sound credible and professional.

Building a Mixing Workflow That Fixes Problems in Order

Many podcasters jump randomly between processing steps, which can create new problems. A better approach is to follow a logical workflow that tackles issues from the most fundamental to the most subtle. This prevents you from, for example, compressing a noisy track and making the background noise more prominent.

  1. Edit first, mix second. Remove coughs, long pauses, and obvious mistakes before applying any processing. This gives you cleaner source material.
  2. Balance levels globally. Use normalization and rough volume automation to bring all tracks to a similar average loudness.
  3. Reduce noise. Apply gates, high-pass filters, and noise reduction plugins to clean up the background.
  4. Fix tonal issues with EQ. Shape the frequency balance of each voice to reduce muddiness, add clarity, and match different microphones.
  5. Control dynamics with compression. Smooth out volume variations after EQ, because EQ can change the apparent loudness of different frequencies.
  6. Deal with sibilance and plosives. De-ess and manually tame plosive bursts.
  7. Final loudness and limiting. Apply a limiter and check integrated loudness (LUFS) to meet platform standards.

Sticking to this order helps ensure each step doesn't undo the previous work. Now let's explore each technique in detail.

Fixing Uneven Volume Levels with Compression and Automation

Uneven volume is the most common complaint in podcast mixing. A host might speak softly while leaning back, then suddenly lean into the mic. Guest speakers often have vastly different projection. The solution combines dynamic processing (compression) with manual adjustments (automation).

Normalization: The First Step

Normalization adjusts the entire clip so its peak reaches a target level, typically -3 dB or -1 dB. This ensures all tracks have similar maximum loudness. However, normalization only sets the ceiling; it doesn't fix internal loudness variations. A quiet passage remains quiet relative to a loud one.

Compression for Smooth, Consistent Speech

A compressor reduces the volume of loud sections and (with makeup gain) raises quiet sections, narrowing the dynamic range. For spoken word, start with these settings:

  • Ratio: 2:1 to 4:1 – gentle enough to avoid sounding squashed.
  • Attack: 1-5 ms – fast enough to catch sudden loud peaks but not so fast that it removes natural transients.
  • Release: 50-100 ms – medium speed so the compressor recovers before the next word.
  • Threshold: Adjust until you see 3-6 dB of gain reduction on the loudest parts of the speech.
  • Makeup gain: Bring the output level back up to match the original perceived loudness.

Listen carefully: over-compression makes speech sound lifeless and fatiguing. If you hear pumping or breathing noises being exaggerated, back off the ratio or threshold.

Volume Automation: The Human Touch

No compressor can perfectly handle every inflection. Manual volume automation lets you ride the level precisely. In your DAW, draw volume points to lower sections where the speaker gets too loud, and raise quiet passages. For example, a guest who speaks softly during an emotional anecdote can be brought forward without affecting the compressor's behavior. Combine automation with compression for the most natural result.

Meeting Platform Loudness Standards

After balancing, measure your mix's loudness using a meter that displays Integrated LUFS (Loudness Units relative to Full Scale). Spotify targets -16 LUFS, Apple Podcasts targets -19 LUFS, and YouTube targets -14 LUFS. Use a limiter or loudness normalization plugin (like Youlean Loudness Meter) to hit your target without clipping. This ensures your episode plays at a consistent volume compared to other content on the platform.

Reducing Background Noise with Gates, Filters, and Spectral Tools

Background noise can be constant (hum, hiss, room tone) or intermittent (keyboard clicks, chair squeaks, page turns). Different noise types require different approaches.

Noise Gates for Silencing Silence

A noise gate mutes the audio when the signal falls below a set threshold. This is excellent for cleaning up low-level noise between words. Set the threshold just above the noise floor (the quietest background sound). Use a fast attack (1-2 ms) so the gate opens quickly when speech begins, and a release of 100-200 ms to avoid a choppy sound. A poorly tuned gate can cut off the ends of words or sound robotic, so listen carefully.

High-Pass and Low-Pass Filters

A high-pass filter (also called a low-cut filter) removes low-frequency rumble below 80-120 Hz. This eliminates most fan hum, air conditioner drones, and structural vibrations. A low-pass filter (high-cut) reduces hiss above 12-15 kHz. Use these as a first line of defense before more aggressive noise reduction.

Noise Reduction Plugins

For constant background noise, use a noise reduction plugin that captures a "noise print" and subtracts it. Tools like Adobe Audition's Adaptive Noise Reduction, iZotope RX's Voice De-noise, or the built-in noise reduction in Audacity can be very effective. Apply conservatively – too much reduction creates an "underwater" or hollow artifact. Aim to reduce noise by 6-12 dB, not eliminate it entirely. For a detailed tutorial, check out Adobe Audition's podcast tutorials.

Spectral Editing for Specific Noises

Advanced DAWs allow you to view audio as a spectrogram (frequency over time). With spectral editing, you can visually select and remove a door slam, a cough, or a dog bark without affecting the speech. This is non-destructive and extremely precise. It's worth learning if you regularly deal with unpredictable background sounds.

Pro tip: The best noise reduction is prevention. Before recording, eliminate sources of noise – turn off fans, close windows, and use quiet keyboards. But when that's not possible, these mixing techniques can salvage the recording.

Preventing Clipping and Distortion with Gain Staging and Limiters

Clipping occurs when the audio signal exceeds 0 dBFS, causing irreversible digital distortion. Distortion can also come from overdriven preamps or microphone capsules. During mixing, you can't fully repair clipped audio, but you can prevent further distortion and, in mild cases, reduce its impact.

Gain Staging Throughout Your Signal Chain

Gain staging means keeping levels healthy at every stage – from the raw track, through each plugin, to the master bus. After any processing (EQ, compression, etc.), check that the output peak doesn't exceed -6 dB. Use trim plugins or channel faders to adjust levels. This headroom ensures that subsequent plugins and the master bus aren't overloaded.

Using a Limiter on the Master Bus

A limiter is a compressor with an infinite ratio that catches and attenuates peaks above a set ceiling. Set the ceiling to -1 dB or -0.5 dB to prevent intersample peaks (peaks that occur between digital samples and can cause distortion on some playback systems). Set the threshold so the limiter only engages on the loudest transients, with no more than 2 dB of gain reduction. Heavy limiting fatigues the ear and reduces dynamics – avoid the "brick wall" approach for podcasts.

Repairing Mild Clipping

If you have a clip with occasional soft distortion (e.g., a peak that just touched 0 dB), you can use a declipping tool. iZotope RX's De-clip reconstructs the lost waveform data. This works best on short, isolated clips. For severe distortion, the best fix is to re-record, but in a pinch you can try spectral repair to mask the artifacts. For more on this, see iZotope's guide to common audio problems.

Enhancing Tonal Quality with Equalization

Inconsistent tonal quality often comes from different recording conditions or microphone characteristics. One voice may sound muddy (too much low-end), another thin (too little low-end). EQ helps balance these differences and bring clarity.

Low-Frequency Roll-Off

Apply a high-pass filter between 80-120 Hz to remove low-end rumble. This cleans up muddiness and reduces the impact of plosives (which we'll address separately). For most voices, avoid boosting below 100 Hz unless the speaker has an exceptionally low voice and you want to add warmth sparingly.

Presence Boost for Clarity

A gentle boost in the 2-5 kHz range adds presence and intelligibility to speech. Be careful – too much boost becomes harsh and brittle. A 1-3 dB shelf or bell filter is usually sufficient. Focus on the 3-4 kHz region for a "forward" sound, which is common in radio broadcasting.

Dealing with Excessive Sibilance via EQ

Sibilance (harsh "s" and "sh") can be reduced with a narrow dip in the 5-8 kHz range. Use a bell filter with a narrow Q (high resonance) and subtract 2-4 dB. This can make your de-esser work less aggressively. However, be careful not to dull the voice – A/B test to ensure you're not removing too much air.

Matching Different Speakers

If one microphone sounds darker and another brighter, try to align their tonal balance. Use a reference track (a section from the brighter mic) to guide your EQ moves. A subtle low-shelf boost on the darker mic and a high-shelf cut on the brighter mic can bring them closer. For more detailed EQ techniques, refer to Sound On Sound's EQ guide.

Controlling Sibilance and Plosives with Dedicated Processors

Sibilance and plosives are frequency-specific problems that can make a podcast sound amateurish if left unchecked. They require separate treatment because they live at opposite ends of the frequency spectrum.

De-essing for Sibilance

A de-esser is a frequency-dependent compressor that reduces gain only when sibilant sounds exceed a threshold. Insert a de-esser plugin after your compressor. Set the frequency to around 7-8 kHz (common for male voices; females might need 8-9 kHz). Adjust the threshold so it triggers only on the harsh esses, not on normal speech. Use a moderate ratio (2:1 to 4:1) and listen in bypass to ensure you're not over-processing. A little reduction goes a long way.

Manual Plosive Taming

Plosives are low-frequency bursts that cause a "pop" sound. While a pop filter during recording is ideal, you can fix them in mixing. Apply a high-pass filter at 80-100 Hz to reduce some of the thump. For stubborn plosives, open the waveform and reduce the gain of the affected region (the first 20-30 ms of a "p" sound) by 6-12 dB. You can also use a spectral editor to remove the subsonic energy while preserving the air sound.

Combining Techniques for Best Results

For the cleanest podcast audio, use a combination: a pop filter during recording, a high-pass filter during mixing, a de-esser for sibilance, and manual editing for the worst plosive bursts. This layered approach prevents any single tool from having to work too hard, preserving natural sound.

For an in-depth look at de-essing, check out Universal Audio's guide to de-essing.

Final Steps: Checking Your Mix and Exporting

After applying the techniques above, do a critical listening pass. Listen on different playback systems: headphones, laptop speakers, and car speakers. Each will reveal different issues. Make small adjustments as needed.

Before exporting, apply a final limiter to catch stray peaks and set your integrated loudness to the target for your distribution platform. Export in a lossless format (like WAV or FLAC) for archival, and convert to MP3 at 128-192 kbps for distribution.

Conclusion

Fixing common podcast audio issues during mixing transforms raw recordings into a comfortable, professional listening experience. By systematically balancing levels, reducing noise, preventing distortion, shaping EQ, and controlling sibilance and plosives, you address the most frequent complaints listeners have. Start with one issue at a time, follow the logical workflow, and always reference your speaker's natural voice. With practice, these techniques become second nature, and your podcast will stand out for its clarity and polish – exactly what your audience deserves.

For further reading on podcast mixing standards and workflows, explore Apple's audio standards for podcasts and Universal Audio's compression guide for spoken word.