Audio Is Half the Story — and Most Editors Treat It as an Afterthought
Watch a video with the sound off and you will judge it on visuals. Watch the same video with headphones on and you will judge it on everything. The pacing lands differently, the jokes hit harder, the emotional beats actually land. That gap is why audio is not a finishing step — it is a structural layer of the story.
The problem is that audio is where production schedules usually break. Scripts get rewritten late, voiceover talent is expensive or unavailable, licensed music costs more than the whole edit, and a two-minute video suddenly needs the same sonic polish as a broadcast spot. Generative audio tools solved the supply problem. They did not solve the taste problem.
This guide walks through a complete workflow for pairing synthesized voice with generated background music: how each layer works, where the failure points are, what settings and levels actually matter, and how to build a process you can repeat every week instead of improvising every time.
How a Modern AI Audio Pipeline Fits Together
Before touching any tool, it helps to see the pipeline as five distinct stages. Each stage has its own quality bar, and problems in early stages cannot be fixed later.
- Script preparation. Clean copy, marked pauses, pronunciation notes, and an emotional beat sheet that maps how intensity rises and falls across the runtime.
- Voice generation. Text-to-speech or voice-cloned narration, rendered per scene rather than as one long take so you can re-roll individual lines.
- Music generation. Mood-matched beds generated with intentional structure — intro, build, loop, outro — and exported as separate stems where possible.
- Mixing. Dialogue intelligibility first, music support second, effects and ambience last. This is where sidechain ducking, EQ carving, and level balancing happen.
- Mastering and delivery. Loudness normalization, true-peak limiting, and a listening check on phone speakers, earbuds, and a laptop.
The temptation is to skip straight to step two and start generating voices. Resist it. A ten-minute script cleanup saves hours of re-rendering because the pronunciation of a brand name was wrong in forty clips.
Text-to-Speech: What Separates a Flat Read From a Performance
Modern neural text-to-speech is genuinely good at the mechanics: clean articulation, natural pacing on well-punctuated sentences, and consistent timbre across a long session. What separates a usable read from a flat one is almost always prosody and direction, not the model.
Prosody, emphasis, and pacing
Prosody is the combination of pitch movement, rhythm, and stress that turns words into meaning. Most synthesis tools expose it indirectly through punctuation, line breaks, and inline emphasis markers. A few practical rules:
- Short sentences read better than long ones. If a sentence has three commas, split it into three lines.
- A paragraph break becomes a longer breath. Use it to create deliberate pauses before a reveal.
- Emphasis markers work best on a single word or a two-word phrase. Marking a whole clause just makes everything louder.
- Speed is a scene-level decision, not a line-level one. Keep a consistent pace within a scene, then shift for the next.
If your tool supports emotion or style tags, treat them as seasoning. One style cue per paragraph is usually enough. Stacking three emotional tags on one line produces the audio equivalent of overacting.
Pronunciation dictionaries and names
Brand names, acronyms, product SKUs, and place names are the number one cause of re-renders. Build a pronunciation sheet before generation begins. For anything ambiguous, spell it phonetically in your notes even if you do not paste it into the tool — it forces you to notice the ambiguity early.
Acronyms deserve special attention. Decide once whether an acronym is spoken letter-by-letter or as a word, and apply that decision everywhere. Inconsistent treatment across a single video sounds like a mistake, because it is one.
Voice cloning, consent, and safety
Cloning a voice is technically trivial now, which makes the ethical question the real constraint. Only clone voices you own or have explicit written permission to use, and keep that permission documented alongside the project files. For commercial work, that means a signed release that covers the specific usage, the duration, and the territory.
There is also a subtler issue: a cloned voice can inherit the vocal habits of its source — filler words, regional vowel shifts, a slightly nasal tone. If your goal is a neutral corporate narrator, a clone of a casual podcast host will fight you the entire way.
Music Generation: From Mood Prompt to Editable Stems
Generated music has one enormous advantage over a traditional library: it can be shaped to the exact runtime of your cut. No more looping a thirty-second track or fading out mid-phrase.
Prompting for structure, not just vibe
Most people prompt with a genre and a mood and accept whatever comes back. Better results come from prompting for structure and function:
- Describe the energy curve: "starts sparse, adds percussion at the midpoint, resolves quietly at the end."
- Name the instrumentation, not just the genre: "felt piano, soft strings, no drums" is far more controllable than "cinematic."
- Specify tempo. Anything 60–80 BPM reads as reflective, 90–110 as neutral narrative, 120+ as energetic.
- State what the track must not do: "no melodic lead, no vocals, nothing in the 1–4 kHz range." That last instruction is your best defense against music that fights the narration.
Stems, loops, and editability
When a tool can export stems — drums, bass, harmony, lead separately — you gain the ability to remove a single element at a key moment. Dropping the drums for four seconds under a critical line is one of the oldest tricks in film sound, and it is trivial when you have stems.
If stems are not available, ask for a version with reduced instrumentation instead and crossfade between the two. A two-second crossfade on a downbeat is almost inaudible and gives you the same effect.
A Repeatable Production Workflow, Stage by Stage
Stage 1: Lock the script and the beat sheet
Number every scene, and for each one write a single line describing the emotional target and whether music should be present at all. Silence is a valid and underused choice. A scene with no music after a scene with music feels like a held breath.
Stage 2: Cast and generate voices
Cast a voice the way you would cast an actor: by register, pace, and texture. If you have more than one character, give each one a clearly different register — one higher, one lower — so listeners can distinguish them without watching the screen. Render per scene, not per project. You will re-roll far fewer lines.
Stage 3: Build the music bed
Generate the full-length bed first, then generate short variations for scene transitions. Keep the bed on a single key and tempo for the whole video unless there is a deliberate tonal shift. Random key changes between scenes are the fastest way to make a project feel assembled rather than composed.
Stage 4: Mix for intelligibility
The dialogue is the product. Everything else serves it.
Stage 5: Master and check on real devices
Normalize to your target loudness, apply a true-peak limiter, and listen on the worst speaker you own. If a line is unintelligible on a phone speaker, fix the mix, not the phone.
Mixing Numbers, Ducking, and Loudness Targets
Exact values depend on your genre, but these starting points work for most narrative and marketing video:
| Element | Typical level | Notes |
|---|---|---|
| Dialogue peaks | −12 to −6 dBFS | Consistent across scenes |
| Music under speech | −22 to −18 dBFS | 12–18 dB below dialogue |
| Music in gaps | −14 to −10 dBFS | Can rise when nothing is spoken |
| Ambience | −30 to −24 dBFS | Barely perceptible, still audible |
| Master loudness | −14 LUFS integrated | Streaming video standard |
| True peak ceiling | −1 dBTP | Prevents codec distortion |
Two techniques do most of the work. Sidechain ducking lowers the music automatically whenever dialogue plays, typically by 6–12 dB with a fast release so the music breathes back in during gaps. EQ carving removes a shallow scoop around 1–4 kHz from the music, exactly where speech consonant clarity lives. Done together, you can keep the music louder than you would expect without sacrificing a single word.
One warning: ducking with an aggressive release creates audible pumping. If you hear the music "breathing" rhythmically, lengthen the release time until it stops.
Spatial Audio and Ambience Without the Mud
Spatial audio is a tool for place, not a decoration. A little reverb tells the listener whether a scene is in a small room or outdoors. Too much tells them the audio is amateur.
Use these cues deliberately:
- Short reverb (0.4–0.8s) for interiors, offices, cars.
- Long reverb (1.5s+) only for halls, canyons, or dream sequences.
- Panning to place a voice slightly left or right of center when the speaker is off-screen or on one side of the frame.
- Room tone as a continuous low-level bed under every interior scene. Cutting reverb tails abruptly between clips is one of the most noticeable amateur errors.
Binaural processing is tempting but risky for video. It is designed for headphones, and viewers on speakers get a phasey, distant mix. If you must use it, check the stereo fold-down before delivery.
Common Mistakes That Undermine Otherwise Good Videos
- Music that never stops. If every second has a bed, nothing feels important. Cut music during the setup so the payoff can land.
- Inconsistent dialogue levels. Scene-to-scene jumps of more than 3 dB read as sloppy even when nobody can name the problem.
- Over-reverbed voice. Generated voices often need less reverb than you think, not more. A dry read with light compression usually sounds more professional.
- Ignoring sibilance. Harsh "s" sounds compound when you add compression. A gentle de-esser early in the chain prevents it.
- Tracks built for the wrong runtime. A three-minute bed cut down to forty seconds loses its structure entirely. Generate at the length you need.
- Skipping the phone check. Half your audience watches on a device with one tiny speaker.
- No negative space. Silence between lines gives listeners time to process, and it costs nothing.
Choosing Tools: Decision Criteria That Hold Up
Tool shopping guides age quickly. Criteria do not. When evaluating any voice or music generator, score it against these:
- Language coverage. Does it handle every language and accent your audience needs, including code-switching within a sentence?
- Direction controls. Can you adjust pace, emphasis, and emotional style per line rather than per project?
- Pronunciation control. Is there a custom lexicon, or must you respell words manually every time?
- Stem or layer export. Can you isolate elements for mixing?
- Licensing clarity. Do you own commercial rights to the output, and are there attribution requirements? Get this in writing before scaling.
- Batch consistency. Does the same voice sound identical across renders on different days?
- API and automation. If you produce more than a handful of videos a month, programmatic rendering pays for itself quickly.
- Watermarking policy. Know exactly what, if anything, is embedded in the output.
Build a small test project — thirty seconds, one voice, one music bed, one transition — and run every candidate tool through it. Demo reels hide weaknesses that a real cut exposes immediately.
FAQ
Will AI narration always sound robotic?
Not necessarily, but it will sound robotic if you feed it long, comma-heavy sentences and ignore pacing. Most of the perceived artificiality comes from uniform rhythm and lack of breath. Shorter lines, deliberate paragraph breaks, and occasional emphasis fix the majority of it.
How loud should music be under a voiceover?
Start 12–18 dB below dialogue peaks and adjust by ear. If you have to concentrate to understand a word, the music is too loud regardless of the number.
Can I use generated music in commercial work?
That depends entirely on the tool's terms. Check the license for commercial use, attribution requirements, and whether the rights transfer to you. Keep a record of the terms as they existed when you generated the track.
How many distinct voices should a video use?
For a short video, one narrator and at most two character voices. Beyond that, listeners lose track unless each voice is dramatically different in register and pace.
Do I need studio headphones?
No, but you need more than one playback system. A pair of closed-back headphones, a phone speaker, and a laptop are enough to catch the problems that matter — mud, sibilance, and level jumps.
Should I mix AI voice with recorded voice?
You can, but match them carefully. Recorded voice carries room tone that synthesized voice does not, so you will need to add matching ambience to the generated lines or remove it from the recorded ones.
What about captions?
Add them. A large share of viewers watch muted, and captions amplify the value of a clean script. If your captions and your voiceover disagree, fix the script — don't let them drift.
How long should I spend on audio relative to editing?
A reasonable rule is one hour of audio work for every three hours of picture editing on a short-form project. It sounds like a lot until you compare it against the re-edit you will otherwise do after the client says the music feels wrong.
Where to Go From Here
The workflow above is intentionally boring: prepare the script, cast a voice, build a bed with structure, mix for clarity, master to a known target, and check the result on bad speakers. Boring is what makes it repeatable.
Start with one small change. Pick a single video you have already published, redo only the audio using these steps, and compare the two versions side by side. The difference will tell you more about your own priorities than any tool comparison ever will — and it will show you exactly which stage of the pipeline deserves your next hour of practice.

