Sound
The Journal

How to de-ess a podcast or spoken voice

By Human Engine Labs · · 5 min read

Sibilance is often worse on spoken word than on singing. Podcast and voiceover mics tend to be bright, they're used up close, and there's no music to mask the harsh esses — so a sharp "s" that you'd barely notice in a song jabs right through a quiet dialogue mix. De-essing a podcast is the same idea as de-essing a vocal, with a few priorities flipped. Here's how to do it so speech stays crisp and easy to listen to for an hour straight.

Key Takeaways

  • Spoken word shows sibilance more than singing — bright mics, close miking, no music to hide it.
  • Intelligibility comes first: tame the harsh esses, never soften them into a lisp.
  • Set it once on the harshest moment and keep it consistent across the whole episode.
  • Low CPU and a live-tracking mode matter when you're processing long takes or recording in real time.

Why podcasts sound more sibilant

Three things stack up against spoken word. Bright, affordable mics — a lot of podcast setups use large-diaphragm condensers or USB mics with a presence lift that exaggerates esses. Close miking — talking a few inches from the capsule maximises the jet of air that makes sibilance (the same reason singers get sibilant up close). And no backing track — in music, a busy mix masks a lot of harshness, but in a bare spoken-word recording the ess has nothing to hide behind. So the exact same voice reads as noticeably harsher on a podcast than on a song.

Intelligibility is the priority

On a podcast, the whole point is that people understand every word, easily, for a long time. That changes how you de-ess:

  • Never over-do it. A lisp on a singer is bad; a lisp across a two-hour podcast is unbearable. Reduce until the esses stop being fatiguing, not until they're gone. (How to spot the line is in how much de-essing is too much.)
  • Aim for "unnoticeable," not "processed." The listener should never think about the esses at all — in either direction.
  • Fatigue is the enemy. Harsh esses are tiring over a long listen even when each one seems minor. Gentle, consistent de-essing across the episode is what keeps it comfortable.

De-essing a podcast, step by step

  1. De-ess on the voice track, early. Before any bright EQ, presence boost, or the loudness processing at the end.
  2. Find the sibilant band. It's usually a touch lower for spoken male voices, higher for brighter ones — the sibilant-range guide covers where to look. Use your de-esser's listen mode to confirm.
  3. Set the amount on the harshest moment in the episode, then leave it. Speech varies, so tune to the worst "s" and let the gentler ones sit.
  4. Keep it consistent. Use the same de-essing across the whole episode (and ideally the whole show) so episodes match and nothing jumps out.
  5. Check on earbuds. Most people listen to podcasts on cheap earbuds and phone speakers, which can exaggerate or hide esses differently than studio monitors — check there too.

Where the right tool helps

Spoken word is where a couple of Sibilance's traits earn their keep. It's phoneme-aware, so it acts on the consonant and leaves the intelligibility-critical parts of speech intact rather than dulling the whole voice. It's light on CPU (measured ~2.5% full-chain), which matters when you're running it across long episodes or a stack of tracks. And it has a low-latency live mode for anyone de-essing while recording or streaming in real time. There's a free tier, which for a lot of podcasters is all they'll ever need.

The short version

Podcasts show sibilance more than songs do — bright mics, close miking, and no music to mask it — so de-ess the voice early, tune it on the harshest moment, and keep it gentle and consistent across the whole episode. Intelligibility and low-fatigue listening come first: tame the esses, never lisp them. A transparent, phoneme-aware de-esser like Sibilance does that without dulling speech, and it's free to start.

Frequently asked

How do I de-ess a podcast?

De-ess the voice track early, before any bright EQ or loudness processing. Find the sibilant band with your de-esser's listen mode, set the amount on the harshest moment in the episode, and keep it gentle and consistent across the whole show so nothing jumps out.

Why does my podcast sound so sibilant?

Three things stack up: podcast mics are often bright, they're used very close to the mouth (which maximises the air that makes sibilance), and there's no music to mask the esses like there is in a song. So the same voice reads as harsher on a podcast than on a track.

How much should I de-ess spoken word?

Less aggressively than you might think. Intelligibility is everything in speech, and a lisp across a long episode is exhausting. Reduce until the esses stop being fatiguing, not until they're gone, and keep it consistent so every episode matches.

Sibilance is the single-authority de-esser this comes from — a real free tier, and a 7-day Pro trial with no card.