Why audio quality decides whether your video gets watched
Audiences forgive a lot: a slightly soft focus, a background that is a little busy, a jump cut that lands a frame late. They do not forgive bad audio. Harsh room echo, a music bed that fights the narration, a voice that sounds like it is reading a weather report — any of these will pull a viewer out of the story faster than any visual flaw. On mobile, where most short-form video is watched, audio is often the only channel carrying the argument, the emotion, and the pacing of the piece.
That is why treating audio as an afterthought is expensive. Retrofitting narration into an edit that is already locked means rewriting sentences to fit shot lengths, re-timing cuts to fit breath pauses, and hunting for music that happens to match a mood you did not design for. The teams that ship consistently good video do the opposite: they plan the audio layer first, generate or record it deliberately, and mix it against a clear target.
AI voiceover and AI music generation have made that planning far more accessible. You can now prototype a narration track in minutes, audition three emotional readings of the same line, and generate an original score that matches the mood arc of your cut — all before a single frame is finalised. The catch is that these tools are only as good as the direction you give them, and direction is a craft. This guide walks through that craft end to end: choosing a synthesis approach, writing scripts that synthetic voices read well, generating music that follows your edit, layering sound design, and mixing everything to broadcast-safe levels.
Plan the audio before you touch the edit
The single highest-leverage habit is writing a short audio plan before editing. It does not need to be long — a page is plenty — but it should lock down five decisions.
Voice persona. Who is speaking? Age range, accent, warmth, energy level, and whether the voice should feel like a friend, a teacher, an analyst, or a narrator. Pick two or three adjectives and stick to them across the whole series so your channel sounds consistent.
Pace and rhythm. A tutorial wants roughly 130–150 words per minute. A dramatic explainer can sit at 110–120. A fast social cut can push 170–190. Pace determines how much script you need per minute of runtime, which in turn determines your production timeline.
Music palette. Two or three genres, a tempo range, and a rule about instrumentation. For example: minimal piano and soft pulse for emotional beats, light electronic percussion for explainers, no drums under dialogue.
Sound design footprint. Decide up front whether you want ambience beds, foley, transition effects, or a deliberately clean sound with nothing but voice and music.
Loudness target. A single number you mix to, every time. More on this below.
Once these are written down, every tool decision becomes easier. You are no longer asking "which voice sounds nice?" but "which voice sounds like the persona we defined, at the pace we specified, sitting under this kind of music?"
Voiceover: choose the right synthesis approach
There are three practical paths, and most projects use a mix of all three.
Stock neural voices
Modern text-to-speech models produce convincing, natural-sounding narration with correct emphasis on most sentences. Stock voices are the fastest option, require no consent paperwork, and are ideal for explainers, product walkthroughs, corporate training, documentary-style narration, and any content where the voice is a vehicle rather than a character.
When auditioning stock voices, do not judge them on a single sample sentence. Paste in a paragraph that contains a question, a list, a number with units, and an emotional beat. Many voices that sound excellent on a greeting fall apart on a five-item list or a rhetorical question.
Cloned and custom voices
Voice cloning lets you create a consistent narrator that matches an existing brand voice or your own voice. The quality is now high enough for serious production work, provided you feed the model clean reference audio: 1–3 minutes of quiet, echo-free speech recorded on a decent microphone, with no music underneath and no overlapping speakers.
Two rules matter here. First, only clone voices you have explicit written permission to use. That means your own voice, a hired performer who has signed a release covering synthetic reproduction, or a client who has authorised it in writing. Second, keep the reference recording on file. If the model drifts or you need to retrain, that file is your source of truth.
Hybrid: synthetic scratch, human final
A workflow many teams overlook: generate the whole narration synthetically as a scratch track, cut the video to it, then re-record the final voiceover with a human performer reading the same script. The synthetic pass acts as a timing blueprint, so the human session takes 20 minutes instead of two hours, and the edit never has to be rebuilt. This is often the best quality-per-hour trade available.
Directing tone, pace, and emotion
Whatever engine you use, the performance is shaped by how you write and label the text. Practical techniques that reliably improve output:
- Punctuation is your primary direction tool. Em dashes create pauses. Ellipses create hesitation. Periods create firm stops. Commas create lift. A sentence with no punctuation will be read as a flat run-on.
- Short sentences carry emphasis. If a line matters, cut it in half.
- Generate line by line for hero moments. For the opening hook and the closing call to action, synthesise each sentence separately and pick the best take. Batch the middle.
- Use pauses deliberately. A 400–700 ms gap before a reveal does more emotional work than any adjective.
- Test two emotional variants. Generate the same paragraph in a warm read and a neutral read, then A/B them with three colleagues. The difference is usually obvious within ten seconds.
Writing scripts that synthetic voices read well
Synthetic narration rewards clean, spoken-language writing. Convert numerals to words when the number is spoken ("fifteen percent," not "15%"). Spell out units on first mention ("kilometres per hour"). Expand acronyms the first time they appear unless the acronym is universally read as a word. Watch homographs — "lead" the metal versus "lead" the verb — and rewrite to remove ambiguity.
Keep sentences between 8 and 20 words. Vary length so the delivery has rhythm. Read every script aloud before you generate it; if you stumble, the model will too. And write contractions ("you'll," "it's") unless you deliberately want formality, because uncontracted speech is the fastest way to make an AI voice sound robotic.
Background music that follows your cut
Music is not wallpaper. It tells the viewer how to feel about what they are looking at, and it does so in the first three seconds. The mistake most creators make is choosing one track for the entire video and fading it in and out. The fix is to design a mood arc.
Scene-aware and emotion-aware generation
Describe your scene the way you would brief a composer: instrumentation, tempo, energy curve, era, and reference mood. "Sparse felt piano, 72 BPM, slowly building strings from 0:45, hopeful but not triumphant, no drums" will produce something far more usable than "emotional background music."
Map the mood arc across your timeline before you generate anything. A five-part structure works for most videos: cold open (curiosity), context (steady), tension (rising), turn (release), resolution (warm and settled). Generate one cue per section rather than searching for one track that does all five jobs badly.
Keep music out of the voice's way
Two technical rules prevent the most common mix disaster. First, avoid tracks with prominent mid-range melodic lines that sit in the same frequency band as speech — solo violin, saxophone, and busy synth leads fight narration. Second, prefer instrumental-only stems for anything under dialogue. If a track has vocals, either remove them or move them to a section with no narration.
Tempo matters more than most people expect. If your edit has a cut every two seconds and your music sits at 140 BPM with a heavy backbeat, the two rhythms will clash. Match tempo to cutting rhythm, or use music with no percussive grid at all.
Ask for edit-friendly structure
When generating music, request formats that serve editing: a 30–60 second seamless loop, stems (drums, bass, melody, pads) exported separately, and a "no melody" alternate mix for use under dialogue. Stems are what let you duck just the melody instead of pulling the whole track down, which keeps energy in the mix while keeping narration clear.
Sound design: ambience, foley, and depth
A clean voice and a good track still sound thin. Depth comes from three layers, added subtly.
Ambience beds. Low-level environmental tone — room hum, distant traffic, wind, café murmur — at roughly −35 to −45 dB relative to dialogue. You should not consciously hear it; you should notice when it disappears.
Foley. Specific object sounds: keyboard clicks, a mug set down, fabric movement, footsteps. Two or three well-placed foley hits per minute add more realism than twenty.
Transition accents. Whooshes, sub-drops, risers, and impacts that mark a cut or a reveal. Use them at structural moments only — a chapter change, a statistic reveal, a reveal shot. If every cut has a whoosh, none of them mean anything.
Sequence layers so they never compete: ambience runs continuously underneath, foley sits at the moment of action, transitions land on the frame of the cut and decay within 400–800 ms.
Mixing for real-world playback
Most viewers watch on phone speakers or laptop speakers in noisy rooms. Mix for that reality, not for studio headphones.
Start with dialogue. Normalise narration so it peaks around −6 dBFS and averages roughly −16 to −18 LUFS integrated across a section. Apply gentle compression at about 3:1 with a slow attack to control peaks without flattening expression. Use a high-pass filter around 80–100 Hz to remove rumble. If sibilance is harsh, a de-esser set sparingly beats a broad EQ cut.
Carve space for the voice in the music. A shallow dip of 2–4 dB between roughly 1 kHz and 4 kHz on the music bus reduces the need for aggressive ducking. High-pass the music bus at 100–150 Hz when the voice is male, or 150–200 Hz when it is higher-pitched, so bass energy does not muddy intelligibility.
Duck the music under speech with a sidechain or a volume automation curve. A 3–6 dB reduction with a 150–250 ms attack and 300–500 ms release sounds natural. Anything deeper sounds like the music is being switched off and on.
Finally, set a delivery target. −14 LUFS integrated is the common standard for general video platforms, −16 LUFS for spoken-word podcast-style audio, and −1 dBTP true peak maximum to avoid clipping after lossy encoding. Check the finished file on a phone speaker before you publish. If it holds up there, it will hold up anywhere.
A step-by-step production workflow
- Write the script and read it aloud. Fix anything you stumble on.
- Draft the audio plan. Persona, pace, music palette, sound design footprint, loudness target.
- Generate narration in sections. Storyboard the video against the scratch voice track rather than the other way around.
- Audition and replace hero lines. Regenerate the hook, the key statistic, and the closing line individually.
- Map the mood arc. Mark the five emotional sections on the timeline.
- Generate or select one cue per section. Request loops and stems where possible.
- Lay in ambience and 2–3 foley accents per minute. Keep them quiet.
- Mix dialogue first, then music, then effects. Duck music under speech.
- Check loudness and true peak. Export at your target, then listen on a phone.
Common mistakes and how to fix them
One track for the whole video. Fix: generate two or three short cues and crossfade at emotional turning points.
Narration written for the eye, not the ear. Fix: shorten sentences, contract words, convert numerals to spoken form.
Music too loud under dialogue. Fix: duck 3–6 dB rather than turning the track down for the entire runtime, and notch conflicting frequencies.
Overloaded sound design. Fix: cut effects until the piece feels empty, then add back one.
Ignoring breath and pause. Fix: insert 300–700 ms gaps before reveals and after key statements.
No consistent loudness across episodes. Fix: mix a reference episode, measure its integrated loudness, and match it every time.
Skipping the phone test. Fix: listen on the worst speaker you own before publishing.
Pre-publish QA checklist
- Narration peaks near −6 dBFS with no clipping and no plosive spikes.
- Integrated loudness matches your delivery target within 1 LU.
- True peak is at or below −1 dBTP.
- Music never masks the narration on a phone speaker.
- No awkward silences longer than 1.5 seconds unless intentional.
- Ambience continues across cuts so scenes do not feel disconnected.
- Every effect serves a structural moment.
- The first three seconds have audio that creates curiosity on its own.
- Captions or subtitles match the final spoken words, not the original script.
FAQ
Do I need a human voice actor at all?
Not for most content. Stock voices handle explainers, training, and narration well. Reserve human performance for cases where character, humour, or emotional nuance is the product itself — and even then, use a synthetic scratch track to lock timing first.
How do I keep an AI voice from sounding flat?
Vary sentence length, use punctuation as direction, generate short lines individually, and add deliberate pauses before reveals. Flatness is usually a scripting problem, not a model problem.
Can I use generated music commercially?
That depends on the specific tool's licence terms. Read them before you publish, keep a record of the licence for each track, and prefer tools that grant broad commercial rights without per-use restrictions.
What loudness should I target for social video?
−14 LUFS integrated with a true peak below −1 dBTP is a safe general target. If you publish to a single platform, check its current recommendation and mix to it consistently.
How long should I spend on audio for a five-minute video?
For a well-planned episode with pre-written script and generated cues, budget 45–90 minutes for generation, editing, and mixing. That includes regenerating hero lines and checking on a phone speaker.
What is the fastest way to improve audio quality?
Fix the narration first. Clear, well-paced dialogue with slightly imperfect music will always beat a beautiful score buried under muddy speech.
Do I need separate stems if I only publish one version?
Yes, for anything you may repurpose. Stems let you build a short social cut, a long-form version, and a silent-background variant without regenerating assets, which pays for itself the first time you republish.



