Guide · 6 min read
Captions and loudness: the two things that decide if anyone watches
Most of the effort in a short video goes into what it says. Most of the drop-off comes from whether it can be read and heard. These are cheaper to fix.
Assume the sound is off
A large share of short video is watched without sound — on public transport, in offices, in bed next to someone asleep. A video that only works with audio loses those viewers in the first second, before your script has a chance.
Captions solve this, but only if they are readable. Three things matter:
- Position. Slightly below the vertical centre. The very bottom of the frame is covered by the interface and the creator handle; dead centre fights with the subject of the image.
- Contrast. A heavy weight with an outline and a soft shadow stays legible over both a dark painting and a bright sky. A caption box works too but costs you a chunk of the frame.
- Pace. Words appearing in time with the narration hold attention better than a full sentence sitting still, because the eye has something to follow. Highlighting the word currently being spoken makes it easier still.
Then assume the sound is on and loud
Loudness is measured in LUFS, and a target between −16 and −14 LUFS puts a short video roughly level with everything else in a feed. Too quiet and viewers reach for the volume — or leave. Too loud and it distorts, which is perceived as low quality rather than as loudness.
Two measurement traps are worth naming. First, peak level is not loudness: a track can peak at maximum and still feel quiet, so measure loudness rather than peaks. Second, if your audio is dual mono — the same signal in both channels — standard loudness meters read it about three decibels louder than it actually is. Measuring a single channel avoids the illusion.
Music sits under the voice, not beside it
Background music should be roughly ten to fifteen decibels below the narration. That sounds like a lot on paper and is barely noticeable in practice; anything closer and the music competes with consonants, which is exactly what makes speech hard to follow on a phone speaker.
One more thing about phone speakers: they reproduce very little below roughly 200 Hz. Deep bass in your music is inaudible to most of your audience while still consuming headroom that could have gone to the voice. Keeping the musical bed above that range makes the mix louder where it counts.
What not to do to a voice
It is tempting to filter a narration track aggressively to make it sound "clean". Be careful with high-pass filters: the fundamental frequency of an adult male voice sits between roughly 85 and 180 Hz, so cutting at 200 Hz removes the body of the voice and leaves it thin and distant. Cut low enough to remove rumble and no higher.