audio-branding-and-storytelling
How to Properly Clean and De-Noise Your Audio Files for Acx Submission
Table of Contents
Why Clean Audio Matters for ACX
The Audiobook Creation Exchange (ACX) enforces strict technical standards to ensure a consistent listener experience across thousands of titles. Submissions that fail to meet these specifications are routinely rejected, costing narrators and producers time and money. Beyond compliance, properly cleaned audio directly impacts listener engagement—background hum, room echo, or plosive distortion all contribute to fatigue and reduced comprehension. Mastering the workflow of de-noising, equalization, and leveling is therefore not optional; it is a core competency for anyone serious about audiobook production.
ACX Audio Specification Primer
Before touching any editing tool, you must internalize the ACX submission requirements. These are non‑negotiable and form the target for every processing decision:
- File format: Monophonic WAV or AIFF, 16‑bit or 24‑bit depth.
- Sample rate: 44.1 kHz (not 48 kHz, not 96 kHz).
- Loudness: Integrated loudness of ‑20 LUFS (±2 LUFS permissible).
- True peak: Maximum peak level of ‑3 dB (no overs).
- Noise floor: Ambient noise must be below ‑60 dB (unweighted).
These parameters are not arbitrary. They guarantee that your audiobook plays back at a consistent volume alongside others on platforms like Audible and iTunes, and that background noise does not mask the narrator’s voice. Always verify your export against these values using a reliable loudness meter (such as the free Youlean Loudness Meter).
Minimize Noise Before the Microphone
The most effective de‑noising happens before you press record. A clean source reduces the processing burden and preserves vocal quality. Consider these pre‑recording steps:
- Treat your recording space with acoustic panels or moveable baffles to reduce reflections and echo.
- Use a quality microphone with a cardioid or hyper‑cardioid pattern to reject off‑axis room sound.
- Maintain a consistent distance from the mic (typically 6–8 inches) to keep the signal‑to‑noise ratio stable.
- Eliminate intermittent noise sources: HVAC hum, computer fans, refrigerator compressors, traffic.
- Record several seconds of “room tone” (silence) during each session. This sample will be used for noise profiling.
Even with a well‑treated room, some level of background hiss or low‑frequency rumble may remain. The following workflow will address those residuals.
Step‑by‑Step Audio Cleaning Workflow
All processing should be applied in a non‑destructive manner (e.g., using clip gain, automation, or FX chains) so you can revert or tweak later. The order matters—applying compression before noise reduction can raise the noise floor and make cleanup harder.
1. Create a Noise Profile and Apply Spectral Reduction
Open your DAW (Audacity, Adobe Audition, Reaper, Pro Tools, etc.) and locate a segment of your recording that contains only background noise—no voice, no mouth clicks. Select 2–3 seconds of this room tone. In Audacity, use Effect > Noise Reduction > Get Noise Profile. In Adobe Audition, use Effects > Noise Reduction / Restoration > Capture Noise Print.
Then, without deselecting, apply the reduction effect. Key parameters to adjust (depending on your software):
- Noise reduction (dB): Start with a moderate value (12–18 dB) to avoid “underwater” or robotic artifacts. Increase only if necessary.
- Sensitivity/Frequency smoothing: Lower values preserve more detail but may leave residual buzz; higher values smooth aggressively. Tune while previewing.
- Attack/release time: Keep these short (around 10–20 ms) so the gate responds quickly to voice onsets.
Audition’s DeNoise module is especially transparent for dialogue. For a deeper cleanup, consider a dedicated tool like iZotope RX Spectral De‑noise, which analyzes both noise and voice in real time and offers advanced control over tonal and broadband components.
2. Remove Low‑Frequency Rumble with High‑Pass Filter
Apply a high‑pass (low‑cut) filter at around 80 Hz with a slope of 12–24 dB/octave. This eliminates traffic rumble, HVAC vibrations, and plosive thumps without affecting vocal clarity. Be careful not to set the cutoff too high (above 120 Hz) as male voices can lose body. Always audition the filter while your narrator is speaking to ensure you aren’t thinning the sound.
3. Equalize for Vocal Presence
After cleaning, apply parametric EQ to shape the frequency content. A common starting point for spoken word:
- Cut below 80 Hz (already done in step 2).
- Reduce 200–400 Hz by 2–4 dB to attenuate boxiness or proximity effect.
- Subtle boost at 3–5 kHz (1–2 dB) to add clarity and intelligibility.
- Gentle shelf above 10 kHz to restore air lost during de‑noising, but avoid boosting hiss.
Use narrow bell cuts for specific resonances (e.g., a room mode at 125 Hz). Listen on good‑quality headphones (like Sony MDR‑7506) and also on consumer earbuds to check for harshness. Over‑EQing introduces phase artifacts and an unnatural timbre.
4. De‑essing for Sibilant Control
Many narrators produce strong sibilance on “s” and “sh” sounds. Insert a de‑esser (or multiband compressor targeting 5–10 kHz) that reduces gain only when those frequencies exceed a threshold. If your DAW lacks a dedicated de‑esser, use a compressor with side‑chain EQ set to that sibilant range. Alternatively, manually edit extreme sibilant peaks with clip gain reduction.
5. Dynamic Range Compression
Compression ensures that quiet passages are not masked by background noise and that loud peaks don’t cause distortion. Use a gentle compression ratio (2:1 or 3:1) with a medium attack (20–30 ms) and fast release (50–100 ms). Set the threshold so that you only catch the loudest 3–6 dB of the dynamic range. Avoid heavy compression that flattens the performance—audiobooks benefit from natural dynamics that convey emotion.
After compression, use clip gain automation to manually match louder and softer sections if your narrator varies widely in volume. This gives you more musical control than relying solely on the compressor.
6. Loudness Normalization and Peak Limiting
The final step before export is to achieve the ACX‑required loudness of ‑20 LUFS with a true peak not exceeding ‑3 dB. Use a loudness meter that measures ITU‑R BS.1770 (many DAWs have integrated such meters, or install a free plugin like TBProAudio mvMeter2).
- Normalize integrated loudness to ‑20 LUFS. This usually involves amplifying the entire file (or segments) uniformly.
- Apply a brickwall limiter at ‑3 dB true peak as a safety measure. Set the ceiling to ‑3.5 dB to allow for inter‑sample peaks. Use a transparent limiter with a look‑ahead of 5–10 ms.
If your recording is already close to ‑20 LUFS and the peaks are under control, you may only need a gentle limiter. Avoid over‑limiting, which creates audible pumping and distortion.
Advanced Cleanup: Clicks, Mouth Noises, and Residual Artifacts
Even after noise reduction, you may notice clicks, pops, or mouth smacks. These can be removed with spectral editing tools:
- Click removal: In Audacity (Effect > Click Removal) or Audition (Effects > DeClicker). Adjust the threshold to detect only the spike while leaving voice untouched.
- Mouth de‑click: Use a dedicated de‑click plugin like iZotope RX De‑click which differentiates between lip smacks and valid speech consonants.
- Spectral repair: If you have a short dropout or a loud transient, use spectral editing (e.g., Audition’s Spectral Frequency Display) to paint over the artifact. This is a last resort—overuse can introduce unnatural textures.
For persistent room echoes or comb‑filtering, consider a de‑reverb plugin (again, iZotope RX or Audition’s DeReverb). Apply it sparingly: too much de‑reverb makes the voice sound hollow.
Exporting for ACX Submission
Once your audio is processed and sounded perfect, export the final mix in the correct format:
- Format: Monophonic (one channel) WAV, 16‑bit or 24‑bit. 24‑bit offers a wider dynamic range and lower noise floor; if your master is clean, 16‑bit is acceptable.
- Sample rate: 44.1 kHz (resample if needed; do this last to avoid quality loss).
- Loudness: Integrated ‑20 LUFS, true peak ≤‑3 dB.
- Metadata: Include title, author, narrator, and track number in the file’s ID3 tags or BWF–iXML (depending on your submission method). Many DAWs allow metadata entry during export.
If you have applied any sample‑rate conversion, use a high‑quality resampling algorithm (e.g., “sinc” in SoX or “medical” in iZotope). Export at the original 44.1 kHz whenever possible to avoid unnecessary processing.
Final Quality Assurance Checklist
Before hitting submit, run a systematic QA pass over the entire recording. Use this checklist:
- Loudness: Measure integrated LUFS across the whole file. It should be within ‑18 to ‑22 LUFS (target ‑20).
- True peak: Scan for any sample that exceeds ‑3.0 dB. If found, adjust limiter ceiling.
- Noise floor: Listen to a silent segment (e.g., 5 seconds of room tone) while the meters are visible. The level should be below ‑60 dB. If not, revisit noise reduction or consider a higher‑quality preamp.
- Audible artifacts: Scan through at moderate volume to catch clicks, strident sibilance, breath noise (excessive breaths can be gated or manually attenuated), and any processing side‑effects.
- Consistency: Compare the first five minutes of the file with the last five—if you used different noise profiles or EQ settings for different chapters, ensure seamless transitions.
Finally, export a short sample (3–5 minutes) and upload it to ACX’s preview player to see if it passes their automated check. Many studios also burn a CD and listen in a car—any flaws will be exaggerated in that environment.
Common Mistakes That Lead to Rejection
Even experienced narrators trip on these pitfalls. Avoid them by understanding why they fail:
Over‑processing and “Washing Out” the Voice
Applying too much noise reduction (e.g., 30+ dB) or heavy‑handed EQ often results in a lifeless, underwater sound. The solution: use the minimum necessary reduction, combine it with subtle EQ boosts in the vocal range, and always A/B against the original to ensure you haven’t lost fidelity. Remember that a tiny amount of broadband noise is more natural than an over‑cleaned void.
Ignoring the Loudness Meter While Editing
Producers sometimes normalize to ‑20 LUFS only at the final export step, unaware that compression and EQ changes had already shifted the loudness. Keep a loudness meter open at all times; it will show you in real time whether your processing is pushing the file toward or away from spec.
Mixing Stereo Files or Using Incorrect Bit Depth
ACX requires monophonic files. If you record in stereo and export as stereo, the submission may be rejected even if the audio is technically clean. Always route your recorded track to a mono output and confirm the export settings. Similarly, submit 16‑bit or 24‑bit—never 32‑bit float, even if your DAW works internally at 32‑bit.
Neglecting to Save Room Tone for Every Session
Noise profiles work best when derived from the actual silence of that session. If you reuse a profile from a different day, you may introduce mismatched frequency content and phase cancellations. Record fresh room tone at the beginning of each recording day, and label it clearly.
Conclusion
Cleaning and de‑noising audio files for ACX is a systematic process that blends technical precision with aesthetic judgment. By understanding the required specifications, setting up a controlled recording environment, and following a logical workflow—noise profiling, rumble removal, EQ, compression, and loudness normalization—you can achieve a polished product that passes automated checks and delights listeners. Always perform a thorough QA pass, and never hesitate to re‑examine any step that introduces audible artifacts. With consistent application of these techniques, your audiobooks will stand out for their clarity and professionalism.
Keep this guide handy whenever you begin a new project, and refer to the ACX official audio submission requirements for the most current specifications. A clean, spec‑compliant recording is not only a requirement—it’s your calling card in the competitive audiobook market.