The Case for Inclusive Podcasting

Podcasting has grown into a mainstream medium, but many creators still overlook a critical audience: people with disabilities. According to the World Health Organization, over 1 billion people live with some form of disability. Without accommodations like descriptive audio and transcripts, your podcast effectively locks out listeners who are blind, have low vision, are Deaf or hard of hearing, or have cognitive processing differences. Accessibility isn’t just about compliance with regulations such as the Web Content Accessibility Guidelines (WCAG)—it expands your reach, improves SEO, and demonstrates a commitment to equity.

The two most impactful accessibility enhancements for audio content are descriptive audio (audio description) and transcripts. Each serves a distinct purpose and requires thoughtful implementation. Below we’ll break down both methods, provide practical editing strategies, and show how to integrate them into your podcast workflow without sacrificing production quality.

Understanding Descriptive Audio

Descriptive audio, also called audio description, is a separate narration track that explains visual information that would otherwise be inaccessible to people who are blind or have low vision. In a podcast context, “visual” elements might include gestures implied by tone, on‑screen actions if the podcast is distributed as video, or even physical objects mentioned but not described. The description is inserted during natural pauses in dialogue so it doesn’t obscure the primary content.

Descriptive audio is not simply reading text on a slide; it’s concise, objective, and timed to fit within moments of silence. For example, if a host says “Yesterday I pulled out my old record player,” the describer might add: “[pauses] the host’s hand moves a vintage turntable onto the table.” The goal is to provide context without interrupting the conversational flow. When done well, listeners who aren’t using the description never notice the inserted track.

When to Use Descriptive Audio

Not every podcast needs constant description. You should add it when:

  • Visual context is essential to understanding – e.g., a video podcast showing a product demonstration.
  • The podcast refers to visual media – e.g., “as you can see in this graph” without verbalizing the data.
  • Physical actions or environments are implied – e.g., a comedy podcast where a host holds up a prop or a true crime podcast where a map is displayed.
  • The podcast is distributed as a video – audio description is a standard requirement for video accessibility under WCAG.
  • The show relies on visual humor or imagery – for instance, a host’s exaggerated facial expression or a visual gag that wouldn’t make sense without description.

For a purely audio podcast, description is less critical because most information is already conveyed through speech. However, if your show includes sound effects or music that carry meaning, you may still need to describe those auditory cues in the transcript (covered below). A good rule of thumb: if you ever say “you can see” or “look at this,” you likely need a description.

How to Add Descriptive Audio: Step by Step

  1. Identify visual elements. Review your episode and note every moment where a listener who cannot see would miss context: gestures, facial expressions, props, on‑screen text, or location changes. Create a time‑coded list.
  2. Draft short, neutral descriptions. Write one or two sentences per visual element. Avoid subjective language (“beautiful sunset”) – stick to observable details (“orange and purple clouds over the ocean”). Keep each description under 10 seconds of reading time.
  3. Time the descriptions. Use your audio editor’s waveform to locate natural pauses or less dense dialogue. Insert your description at a point where it doesn’t overlap with important speech. If no pause exists, you may need to adjust the original audio slightly – shortening a breath or trimming a silence. Never cut into words; extend a pause instead.
  4. Record the narration. Use a clear, neutral voice – ideally the same voice throughout. If you use a separate audio track, you can adjust its volume and timing independently. For consistency, record all descriptions in one session.
  5. Edit the description track. In your DAW (e.g., Audacity, Adobe Audition, Reaper), place the description clips on a secondary track. Adjust fades to ensure a smooth transition. Listen with headphones to confirm the description does not clash with dialogue. Use subtle volume ducking (1–3 dB) on the main track during the description if needed.
  6. Test with real users. Recruit a few people who are blind or have low vision to listen to the described version. Ask them if anything was missed or if descriptions felt intrusive. Iterate based on feedback. Also test with sighted listeners to ensure the descriptions don’t break the flow.

For advanced workflows, consider using the W3C’s Audio Description specification as a technical reference. Many media players now support multiple audio tracks, allowing listeners to toggle descriptions on/off. Some podcast apps like Overcast also support chapters, where you can embed descriptions as separate tracks.

Common Pitfalls in Descriptive Audio

  • Over‑describing: Not everything needs description. If a gesture is irrelevant to the content, skip it.
  • Describing during dialogue: Always place descriptions in pauses. If there’s no pause, rewrite the host’s lines to create a brief gap.
  • Using subjective language: Stick to observable facts. Instead of “sad face,” say “host’s eyebrows furrow and corners of mouth turn down.”
  • Inconsistent voice: Use the same narrator for all descriptions in an episode and ideally across your whole show.
  • Ignoring timing: A description that starts too late or ends too soon can confuse listeners. Use your DAW’s time‑stretching to adjust speech rate only as a last resort.

Tools for Descriptive Audio

  • Audacity (free) – supports multiple tracks, silence detection, and fade effects. Best for beginners. Use the Label Tracks feature to mark description points.
  • Descript – AI‑powered editor that lets you type to edit audio; useful for inserting text that becomes synthesized speech, though you may still want human narration for quality. It also can generate a transcript that you can edit alongside the waveform.
  • Adobe Audition – professional multitrack editing with spectral display to pinpoint pauses. Includes a “Treat as Mono” option and clip‑based effects for quick EQ matching.
  • Auphonic – not for editing, but its leveler can normalize the description track to match the main audio. Use it as a final step to blend the two tracks.
  • Reaper – affordable, highly customizable DAW with excellent MIDI and marker functionality. Ideal for complex accessibility projects.

Providing Transcripts

A transcript is a text version of everything that is said in a podcast. Transcripts are the single most effective accessibility feature for a wide range of users: Deaf or hard‑of‑hearing listeners, non‑native speakers, people with auditory processing disorders, and anyone in a noisy environment. They also boost search engine optimization because search engines can index the text, making your episode discoverable via quotes or topics. In fact, podcasts with transcripts consistently rank higher for long‑tail keyword searches.

There are two common types of transcripts:

  • Basic verbatim transcript – every word, including filler words (um, ah), false starts, and stutters. Useful for research or legal accuracy but can be hard to read.
  • Edited transcript – cleaned up for readability while preserving meaning. Removes fillers, corrects grammar, and may add paragraph breaks. This is preferred for most podcasts.

Whichever style you choose, you must include speaker identification, timestamps at regular intervals (every 5–15 minutes), and descriptions of non‑speech audio like sound effects, music, or laughter. For example: [door creaks], [upbeat synth music], [audience laughs]. For music, note the title and artist if relevant.

Creating Effective Transcripts: Best Practices

  1. Transcribe accurately. Use automated speech recognition (ASR) as a starting point, then manually proofread. Services like Otter.ai, Rev.com, or Descript offer AI‑generated drafts. Plan for 3–4 times the episode length in editing time for a full manual pass.
  2. Identify speakers. Use labels like “Host:” or “Guest:”. If there are multiple guests, use names or roles (e.g., “Dr. Smith:”). Avoid generic labels like “Speaker 1” when the identity is known.
  3. Add timestamps. Mark time codes at natural breaks – every 5 minutes is a good rule. Format as [00:05:00]. Some podcast apps support clickable timestamps that jump to that moment in the audio.
  4. Describe non‑speech audio. Write cues in brackets: [suspenseful music], [phone rings], [audience laughter]. For music, note the title and artist if relevant. Avoid generic descriptions like “music playing” – specify the mood if it adds context.
  5. Format for readability. Use paragraphs, avoid blocks of text longer than 4‑5 lines. Include bold or italics for emphasis if needed. Use headings for topic changes if the transcript is long.
  6. Publish alongside the audio. Link to the transcript from the episode page, podcast app notes, and show notes. Use a separate page or a collapsible section on the same page. Consider providing the transcript as a downloadable PDF or an HTML page. HTML is preferred for screen reader compatibility.
  7. Keep it accessible. Use a clean font, sufficient contrast, and avoid PDFs that are not tagged. HTML transcripts are the most accessible format because they work with screen readers. Add a skip‑to‑transcript link at the top of your episode page.

For video podcasts, you also need captions synchronized with the audio. While transcripts are the raw text, captions are time‑coded and displayed on screen. Many platforms (YouTube, Vimeo) auto‑generate captions, but you should review and correct them. The FCC has regulations for television, but for digital‑first podcasts, WCAG Level A requires captions for all prerecorded video content. If you publish on YouTube, use their caption editor or upload a properly timed .SRT file.

Tools for Creating Podcast Transcripts

  • Otter.ai – AI transcription with speaker identification; exports in plain text, SRT, or VTT for captions. Free tier includes 300 minutes per month.
  • Descript – transcribes audio automatically and lets you edit audio by editing the text; also exports transcripts with timestamps. The text‑based editing feature makes it easy to clean up filler words.
  • Rev.com – human‑powered transcription with high accuracy (98%+) for a fee. Ideal for content where precision is critical, such as legal or medical podcasts.
  • OpenAI Whisper – free, open‑source model you can run locally or via API; supports many languages. Good for budget‑conscious creators who have some technical skill.
  • Sonix – automated transcription with collaborative editing features. Includes a browser‑based editor where you can correct timestamps and speaker labels.
  • Temi – low‑cost AI transcription with decent accuracy (80–90%). Best for quick drafts that you then manually edit.

Integrating Accessibility into Your Production Workflow

Rather than treating descriptive audio and transcripts as afterthoughts, bake them into your podcast editing pipeline. This saves time and ensures consistency. Below is a phased approach that weaves accessibility into every stage of production.

Phase 1: Pre‑production

  • Plan segments that may need description (e.g., visual demos, slides).
  • Write a rough script that includes natural pauses for description. Mark these pauses in your script with brackets.
  • Inform guests that you will be adding accessibility enhancements; they may need to describe visual cues verbally during recording (for example, “I’m pointing to the third row of the chart”).
  • Decide whether you will produce a verbatim or edited transcript, and communicate that to any transcription service.

Phase 2: Editing

  • Record the main audio and the description track in separate channels. If you record descriptions in post, use a marker system in your DAW to flag moments requiring description during editing.
  • Use a marker system in your DAW to flag moments requiring description.
  • Transcribe the clean edit of the final mix before adding descriptions (so transcript matches the final published audio). This avoids mismatch between transcript and audio.
  • While editing, add description track clips and adjust timing. Use crossfades of 20–50ms to avoid clicks when descriptions start/end.
  • Export a mono mix of the main audio for processing with Auphonic to level the loudness (‑16 LUFS for podcasts). Then blend the description track at a slightly lower level (‑18 LUFS) so it doesn’t dominate.

Phase 3: Publishing

  • Host the podcast on a platform that supports multiple audio tracks (e.g., Podcast.co, or use chapters with external description files).
  • Embed the transcript in the episode show notes or on a dedicated page. Use the <details> HTML element to keep it from cluttering the page but still accessible. Add a direct “Download transcript” link.
  • Add a link to download the transcript and the descriptive audio file. If you use separate tracks, provide instructions for listeners on how to switch tracks in their app.
  • Test the user experience with any available accessibility tools (e.g., Windows Narrator, macOS VoiceOver). Verify that the transcript is navigable by heading and that description files play correctly.
  • Add an accessibility statement to your podcast website describing what features are available and how to request accommodations.

Remember that accessibility is not a one‑time checkbox. As you produce new episodes, maintain a style guide for descriptions and transcripts. Regularly solicit feedback from listeners with disabilities and update your practices accordingly. Consider setting a quarterly review of your accessibility features with a small user test group.

Depending on where you operate, accessibility may be legally mandated. The Americans with Disabilities Act (ADA) and Section 508 in the United States, the Accessibility for Ontarians with Disabilities Act (AODA) in Canada, and the European Accessibility Act (EAA) all cover digital content. While smaller independent podcasters may not be directly targeted, large media organizations, universities, and government agencies are required to comply. The EAA, which took effect in 2025, applies to any digital product sold in the EU, including podcasts offered as part of a commercial service.

Beyond legalities, consider the potential of accessible podcasts to strengthen your community. Listeners who rely on transcripts often become loyal followers because they feel included. They are also more likely to share your content on social media, where transcripts make it easy to quote key insights. In a recent survey, 70% of disabled podcast listeners said they would stop following a show that lacked accessibility features. Including these features can also reduce your liability risk and improve your reputation among advocacy groups.

Measuring the Impact of Accessibility

Once you implement descriptive audio and transcripts, track how they affect your podcast metrics. Look for:

  • Increased organic traffic from search engines after transcript pages are indexed.
  • Higher listener retention among users who prefer reading along.
  • Positive feedback from listeners with disabilities, especially through direct messages or reviews.
  • Growth in shareability – transcripts make it easier to share quotable excerpts and create social media snippets.
  • Reduction in support requests asking for alternative formats, if you previously had none.

Use tools like Google Analytics to monitor page views on transcript pages. Compare episode downloads before and after adding transcripts. If you use a dynamic ad insertion platform, note that transcript content can also improve ad targeting.

Conclusion

Adding descriptive audio and transcripts to your podcast is not only a technical edit – it’s a commitment to universal design. Descriptive audio opens your show to people who are blind or have low vision, while transcripts serve everyone from the Deaf community to multitaskers to non‑native speakers. The tools and workflows exist today to implement both without requiring a massive production overhaul. Start with one episode, gather feedback, and iterate. Your podcast will be richer, more discoverable, and more inclusive because of it.

To go deeper, explore the W3C Web Accessibility Initiative (WAI) resources on making audio and video media accessible. For community support, join online groups like the Podcast Accessibility Network or follow podcasters who already model inclusive practices. The NPR transcription standards also offer a professional benchmark for quality and consistency.