Sound
The Journal
De-essingVocalsMixing

How to de-ess vocals without dulling them

By Human Engine Labs · · Updated · 7 min read

We have a soft spot for the humble de-esser. It does one unglamorous job — take the harsh ess out of a vocal — and almost every de-esser we've ever reached for does it a little wrong. They tame the harshness by dimming the whole top end, and you end up with a vocal that's smooth, sure, but also kind of… asleep. No air. No sparkle.

It doesn't have to be that trade. Here's how we think about removing only the sibilance and leaving everything you actually like about the voice alone.

Sibilance isn't a frequency

This is where most de-essers go wrong on the first move. Sibilance isn't a spot on the spectrum you dial in once and forget. It's a set of sounds — the consonants you make with a little jet of air: s, sh, z, ch, t, and the harder f and th.

Their energy usually lands somewhere between 5 kHz and 10 kHz, but "usually" is doing a lot of work in that sentence. A bright voice might spike up at 8 or 9 kHz; a deeper one closer to 5 or 6. It moves with the singer, the mic, even the word.

So if you park a static dip at 7 kHz and call it de-essing, you'll miss every ess that lands somewhere else and carve a hole in the ones that don't. The target keeps moving. Your de-esser has to move with it.

Why the usual approach dulls the voice

A classic de-esser is just a compressor listening to a high band. When anything up there gets loud, it pulls the whole band down. Trouble is, a brightly sung vowel, a breath, and a cymbal bleeding into the vocal mic all live up there too. So it fires on things that aren't esses — and every time it fires, it takes the air with it.

Push it hard enough to catch the worst esses and you've basically built a top-end ducker that flattens the vocal every few words. That's the dullness. It's not that you removed too much sibilance; it's that you removed a lot of things that weren't sibilance at all.

Remove the ess, not the octave

Three things separate transparent de-essing from top-end ducking:

  • Follow the phoneme, not a fixed band. Act on where the sibilance actually is on this word — not on a frequency you guessed in advance.
  • Split narrowly. Reach for a tight region around the ess and the air survives. One broad split is why so many de-essers sound dull.
  • Only duck when it's really an ess. A bright vowel is not an ess. A de-esser that can tell the difference simply won't touch the vowel — that's the whole ballgame between a level trigger and one that actually classifies the consonant.

This is, honestly, the itch we built Sibilance to scratch: a detector that recognises the consonant and protects it, splitting narrowly so it reduces the ess and leaves the vowel's high end where it was. On our test material that comes out to about 5 dB of sibilance gone while the non-sibilant highs stay within ±0.04 dB — and we show our working, because "trust us" isn't a measurement.

A starting point that works

  1. Find the ess. Solo the vocal and listen for the words that bite. If you can watch a spectrum, the esses flash as little bursts up top.
  2. Set the amount on the worst word, by ear. Pull sibilance down until the harshness leaves the hardest ess — then stop. If the quieter esses start to lisp, you've gone past it.
  3. Check the vowels. Long, open, sung vowels should be untouched. If they dip when the singer really opens up, your de-esser is chasing brightness, not sibilance.
  4. Listen to what you're removing. This is the one that changes everything: solo just the signal the de-esser is taking out. You want to hear esses and almost nothing else. If you hear whole words, or tone, or breath — ease off. (We call this Δ Listen; most plugins have a "listen" or "diff" mode. Use it. Every time.)

Please don't reach for a static EQ

We get the temptation — a permanent high-shelf cut or a notch to "fix" a harsh vocal. But sibilance is a flicker. It's there for a few milliseconds per word and gone the rest of the time. A static cut is dull all the time to solve a problem that only shows up some of the time. That's the whole reason de-essing is dynamic: it acts when the ess is happening and gets out of the way the instant it's over.

Get those two halves right — act only on the ess, and only while it's happening — and the vocal comes out both smooth and open. That's the entire trick. There isn't a secret second one.

Sibilance is the single-authority de-esser this comes from — a real free tier, and a 7-day Pro trial with no card.