Why audio decides whether a video works
A viewer will forgive soft focus, a slightly crooked horizon, or a plain background. Almost nobody forgives bad sound. When narration is muffled, when music fights the dialogue, or when the mix jumps between loud and quiet scenes, the audience leaves — often within the first fifteen seconds — and no amount of color grading brings them back. Audio is the fastest signal of production quality, because the ear processes it faster than the eye processes the frame.
That is why an audio-first mindset pays off in every video workflow. It changes what you build first, how you write, and what you check before export. It also changes which tools you need. Modern speech synthesis can carry tone, hesitation, emphasis, and breath. Modern generative music can deliver stems, tempo, and structure that follow an edit instead of fighting it. The models are no longer the bottleneck. The workflow is.
The sections below lay out a complete, tool-agnostic approach: how to plan the sound of a video, how to produce narration and score, how to mix to delivery standards, and how to avoid the mistakes that make synthetic audio sound synthetic.
The four audio layers every video needs
Treat video sound as four separate layers. Each has a different job and a different set of rules.
- Voice layer. Narration, interviews, dialogue, or character performance. This is the layer that carries information and personality. It must be intelligible before it is beautiful.
- Music layer. Beds, transitions, and emotional cues. Music sets pace and tells the viewer how to feel about a shot. It should support, not compete.
- Ambience layer. Room tone, city hum, wind, crowd murmur. This is the glue. Silence in a video feels like a technical error; a quiet ambience bed makes cuts feel intentional.
- Accent layer. Door slams, whooshes, keyboard clicks, risers. These are punctuation marks. A handful used well is worth more than a hundred used randomly.
| Layer | Priority | Typical level | Main risk |
|---|---|---|---|
| Voice | Highest | -6 to -3 dBFS peaks | Noise, clipping, robotic delivery |
| Music | High | -18 to -14 dBFS under speech | Masking dialogue |
| Ambience | Medium | -30 to -24 dBFS | Distracting loop points |
| Accents | Low | -12 to -8 dBFS peaks | Overuse, sameness |
Plan all four before you generate anything. A video whose voice is excellent but whose ambience is missing still sounds unfinished.
Voiceover workflow: from script to finished narration
Write for the ear, not the page
Most robotic-sounding narration was robotic on paper before it was ever synthesized. Write in short sentences. Use contractions. Put one idea in each sentence. Spell out numbers, dates, and symbols unless the voice engine handles them reliably, and then check every one of them anyway. Mark the beats you want: a comma for a short breath, a period for a full stop, an em dash for a dramatic pause.
Read the script out loud, then cut every sentence you stumbled over. If you cannot read it smoothly, a model will not either.
Directing a synthetic voice
Voice selection is the highest-leverage decision in the whole process. Judge candidates on register (low, mid, bright), pace, warmth, accent, and how they handle questions and lists. Do not choose from a demo reel of a single cheerful sentence. Build a test paragraph that includes:
- A neutral statement of fact
- A question
- A three-item list
- A number and a proper noun
- One line that needs genuine emotion
Generate that paragraph with every candidate, listen on headphones and on a phone speaker, and pick the voice that survives both.
Then use the fine controls. Stability or consistency settings keep a long read from drifting. Style or emotion settings push delivery toward a mood. The most common error is pushing style too hard — exaggerated settings produce a performance that feels like a caricature. Aim for a slight lift, not a leap.
Pronunciation, pacing, and pickups
Build a pronunciation dictionary for brand names, technical terms, and acronyms. Update it as you go; it becomes an asset for the entire series.
Split long scripts into blocks of 30 to 90 seconds and generate them separately. When one line changes, regenerate that block only. Keep a simple change log so you know which block maps to which section of the timeline.
Finally, listen at 1.25x speed. Problems with pacing, sibilance, and misplaced emphasis become obvious when compressed. Fix them before you move on.
When a human voice is still the right call
Synthetic narration is excellent for explainers, tutorials, product walkthroughs, internal training, and localized versions of existing content. Consider a human performer when the piece depends on comedic timing, fine emotional nuance, legally sensitive statements, or a personal brand built on the speaker's identity. A useful compromise: use a human for the signature intro and a synthetic voice for the bulk of the content.
Music workflow: matching score to the edit
Lock the picture first
Scoring before picture lock is the classic way to waste a day. Music written for a cut that later loses six seconds will land wrong on every beat. Lock the edit, export a reference file, and only then decide where music belongs. Use a temporary track while you edit — something you will absolutely not publish — and note the exact timecodes where it works.
Prompt for structure, not just mood
A prompt like "sad piano" produces generic results. Describe what the cue must do instead:
- Instrumentation and texture
- Tempo in beats per minute
- Key or tonal center, if you care about that
- Energy curve: where it builds, where it drops out, where it resolves
- Length and whether it needs a clean ending or a loopable tail
- What should be absent — for example, no drums, no vocals, no high strings
Reference broad genres and eras rather than specific copyrighted tracks. "Warm analog synth with a slow pulse" is both safer and more useful than naming a song.
Build a cue sheet
A cue sheet is a table you fill in before generating anything. It keeps music purposeful and prevents the temptation to lay one long track under the whole video.
| In | Out | Function | Energy | Notes |
|---|---|---|---|---|
| 00:00 | 00:12 | Cold open bed | Low | Dialogue over it, no drums |
| 00:12 | 00:48 | Main theme | Medium | Builds to the logo sting |
| 00:48 | 02:10 | Interview support | Very low | Drops out under key quotes |
| 02:10 | 02:40 | Climax | High | Full mix, resolves near the end |
One video usually needs four to six cues, not one continuous track. Short, purposeful cues are easier to mix, easier to replace, and easier to re-use in future videos.
Use stems and cut to the beat
Ask for stems — drums, bass, harmony, melody, ambience — instead of a single stereo file. With stems you can mute the percussion under a quiet passage, extend an intro by looping the harmony, or fade only the melody at the end. This is the difference between music that fits and music that merely plays.
When you cut, align transitions to musical beats. A cut landed two frames before the downbeat feels accidental; landed on the downbeat, it feels designed.
Ducking and dialogue priority
Music exists to support the voice. Set music beds roughly 14 to 22 dB below narration during spoken passages, and lift them in the gaps. Use volume automation or sidechain compression rather than a static level, because a constant bed either buries the voice or vanishes when nobody is talking.
End-to-end production workflow in eight steps
- Lock the edit and export a reference cut. Include timecode and total duration.
- Map the audio. Write the cue sheet, note every narration section, and mark where ambience changes.
- Draft the script and a scratch read. Record yourself on a phone. It reveals awkward phrasing instantly.
- Finalize the script and pronunciation list. Freeze it. Late script changes are the main cause of remixes.
- Generate narration in blocks. Check each block against the script, verifying numbers, names, and emphasis.
- Generate music cues. Export stems for each cue and name files by timecode and function.
- Edit the audio. Clean noise, remove distracting breaths, EQ for clarity, de-ess, and level each layer.
- Mix, normalize, and export. Hit your loudness target, verify true peak, and deliver the formats your platform needs.
If you work alone, do steps three through six in one sitting and mix on a different day. Fresh ears catch what tired ears forgive.
Loudness, mixing, and delivery specs
Mixing to a target is not optional. Platforms normalize playback, so a mix that is too loud gets turned down and a mix that is too quiet gets turned up — along with its noise floor. Common targets:
- Online video platforms: around -14 LUFS integrated, true peak no higher than -1 dBTP
- Podcast distribution: around -16 LUFS integrated, -1 dBTP
- Broadcast in North America: -24 LKFS
- Broadcast in Europe: -23 LUFS
Export at 48 kHz, 24-bit for video work. Check mono compatibility, because a surprising number of viewers watch on a single phone speaker. Test the final mix on headphones, a laptop speaker, and in a car, in that order.
Keep a little headroom. A voice that peaks at -3 dBFS leaves room for music and effects without clipping later.
Rights, licensing, and provenance
Before publishing, answer four questions in writing:
- Do you have the right to use this voice? If a voice was cloned, is there documented consent? Does the tool's terms of service permit commercial use and monetization?
- Do you have the right to use this music? Some generated music is unrestricted, some requires attribution, and some prohibits certain uses. Save the terms you agreed to.
- Can you prove what you made? Keep project files, prompts, generation dates, and raw outputs. Provenance records settle disputes quickly.
- Do you need to disclose it? Many platforms and several jurisdictions require labeling of synthetic or altered media. A short line in the description is cheap insurance.
If a voice or a track is central to your brand, consider commissioning a human and licensing the result outright.
Common mistakes that make AI audio feel artificial
- One giant script block. Long generations drift in tone, pace, and volume. Use blocks.
- Over-styling the performance. Pushing emotion sliders to maximum produces melodrama.
- Ignoring silence. A half second of nothing before a key line is a technique, not a mistake.
- Wall-to-wall music. Constant score flattens the emotional curve. Let cues end.
- No pronunciation dictionary. Wrong names destroy credibility faster than any other detail.
- Mixing on one system. Check on at least three playback devices.
- Skipping loudness normalization. Quiet exports get boosted along with their hiss.
- Mismatched accents and audience. A voice that clashes with audience expectations creates friction before your message lands.
- No rights check. The cheapest-sounding problem is the one that gets your video taken down.
Choosing tools: decision criteria
Rather than chasing the newest model, evaluate tools against your actual production needs. Useful criteria:
- Prosody quality. Does the voice handle questions, lists, and emphasis naturally?
- Control granularity. Can you control pace, pitch, pauses, and per-word emphasis?
- Languages and accents. Does it cover every market you publish in, and can you check those languages yourself?
- Export options. Stems, WAV, 48 kHz, and clean tails matter more than a flashy interface.
- Automation. A command-line or API path lets you regenerate hundreds of lines consistently.
- Licensing clarity. Written commercial terms beat optimistic marketing copy.
- Pricing model. Usage-based pricing suits bursts of work; flat subscriptions suit steady volume. Match the model to your cadence.
- Editor integration. Audio that lives next to your timeline saves hours of file wrangling.
A practical stack usually has four parts: an all-in-one editor for assembling and rough mixing, a specialist speech synthesis tool for narration, a music generator for cues and stems, and a restoration tool for noise, hum, and reverb cleanup. Nothing needs to be expensive. It needs to be consistent.
FAQ
Can synthetic narration replace all human voice-over?
For informational and instructional content, often yes. For performance-driven content — comedy, emotional storytelling, brand signature — a human still wins. The practical answer is a hybrid: automate the volume, hire for the moments that carry meaning.
How long should a music cue be?
As long as the section it supports, no longer. Most cues land between 15 and 90 seconds. If a cue runs the length of the video, break it into sections with distinct energy.
Do I need stems?
If you plan to mix dialogue on top of music, yes. Stems let you remove drums under quiet passages, extend intros, and fix endings without regenerating the whole track.
How do I keep a voice consistent across a series?
Freeze the voice, the settings, and the pronunciation dictionary. Save them as a preset and always generate from the same template. Consistency comes from process, not from luck.
What export settings should I use?
48 kHz, 24-bit WAV for audio you will still mix. Deliver compressed audio only at the final step.
Is generated music safe to monetize?
It depends on the terms you agreed to. Read them, keep a copy, and avoid naming existing songs in prompts. When in doubt, use a tool whose license explicitly permits commercial use, or commission a composer.
How much time should mixing take?
Budget 30 to 60 percent of your post-production time on audio. It is not the last five minutes of the project; it is half the project.
What if the voice mispronounces a word?
Add it to the dictionary, regenerate only that block, and drop the new take into the timeline. Do not stretch audio to fake it.
A short pre-export checklist
Before you export, confirm that narration is clear on a phone speaker, that music never masks a word, that ambience runs underneath every cut, that loudness hits your target and true peak stays under -1 dBTP, that files are named by timecode and function, that rights for the voice and music are documented, and that synthetic media is disclosed where required.
Do that, and the audio stops being the thing you apologize for. It becomes the reason people watch to the end.



