top of page

Why Bad Audio Makes Good Video Look Cheap

8 minutes ago
3 min read

There is a moment in a lot of first round reviews where the client says something feels off, and then starts talking about the picture. The grade, maybe. The framing.

Often the picture is fine. What they are reacting to is a room with hard walls, recorded at eleven in the morning, with a faint hum from the air conditioning that nobody noticed on the day.

Video audio quality is the most reliable predictor of whether a piece of content reads as professional, and it is also the line item people cut first.

The research on this is fairly blunt

In 2018, Eryn Newman and Norbert Schwarz ran a study where people watched the same scientist deliver the same talk, using the same words, in two different audio qualities. With poor audio, viewers rated the research as worse, the scientist as less intelligent, and the work as less important. You can read the paper in Science Communication, or USC's plain English write up of it.

Nothing about the content changed. Only the sound did.

The mechanism is processing fluency. When something is harder to take in, people attribute the difficulty to the thing itself rather than to the conditions they are taking it in under. A voice that is tiring to listen to starts to feel like a person less worth listening to, and the viewer never consciously registers the swap.

The room matters more than the microphone

The fix most people reach for is a better microphone. The problem is usually the space they are standing in.

A modest mic in a treated room will beat an expensive one in a glass meeting room every time. Hard parallel surfaces bounce sound back into the mic a few milliseconds behind the direct signal, and that smearing is what makes a recording sound amateur even when everything else has been done right.

You can do a great deal with very little. Soft furnishings, a rug, curtains, bodies in the room, a duvet on a stand just outside frame. None of it photographs well and all of it works.

Five things we check on every mix

  • Noise floor: the hum, the traffic, the fridge, the air conditioning. Some of it can be removed later and some of it cannot, which is why we listen for it on the day rather than in the edit.

  • Levels: one speaker should not be noticeably louder than the next, and a viewer should never have to reach for the volume control halfway through.

  • Plosives and sibilance: the hard p sounds and sharp s sounds that survive everything else and make three minutes feel like ten.

  • Music under dialogue, which is where most mixes come apart.

  • The phone test: the final mix gets checked on a phone speaker, because a phone speaker is where it will actually be heard.

The music problem

An editor spends four hours with a track. By hour four they have stopped hearing it, so they push it up, because to them it has receded into the background.

To a viewer hearing it for the first time it is now sitting on top of the dialogue. This is the single most common note we give on work that arrives with us from elsewhere.

Our habit is to set the music level at the very end of the process, on fresh ears, on a phone, and then pull it down a little further than feels right.

Sound, look, say

We describe our craft in three parts: sound, look and say. Sound is listed first on purpose.

A viewer will forgive a slightly soft image, an awkward cut, a grade that is not quite there. They will not forgive audio they have to work at. They leave, and they could not tell you why if you asked them.

The takeaway

If you have a budget decision to make between a better camera and a better sound setup, take the sound. It costs less, it is more noticeable, and it changes whether people believe the person on screen.

More on how we handle this across our work.

 
 
 

Comments


Commenting on this post isn't available anymore. Contact the site owner for more info.
bottom of page