Audio Is Half the Story, and It's Usually the Rushed Half
Viewers forgive a lot. A slightly soft focus, a jump cut that lands a beat late, a title card with a typo — most of that slides by. What they do not forgive is bad audio. Muffled narration, a music bed that fights the voice, an abrupt level jump between scenes: these are the things that make people tap away within eight seconds, often without ever being able to say why they left.
That asymmetry is the reason AI voice and background music generation deserve more than a five-minute slot at the end of your edit. Synthesized narration has quietly crossed the threshold from "obviously robotic" to "plainly competent," and it improves every few months. But the models only give you raw material. The difference between a video that sounds professionally finished and one that sounds like a demo comes down to direction, timing, and a small number of post-production decisions that almost nobody teaches.
This guide is a practical workflow for building both halves of the soundtrack — narration and score — and then getting them to cooperate. It covers prompt design for voice, emotional mapping for music, sync technique, loudness standards, and the mistakes that show up again and again in AI-assisted productions.
The Two-Track Mindset: Narration and Score Are Separate Systems
The single biggest structural error people make is treating audio as one task. It is not. Narration and music obey different rules and should be generated, edited, and judged separately before they ever meet in a timeline.
Narration is information. Its job is comprehension. Every decision — pace, articulation, room tone, level — should be evaluated by asking whether a listener on a phone speaker in a noisy room can follow it.
Music is emotion. Its job is to tell the viewer how to feel about what they are seeing. It can be vague, repetitive, and texturally simple. If a viewer consciously notices the music, it is usually doing too much.
Because their goals differ, their quality checks differ. Narration gets checked for clarity and consistency; music gets checked for arc and restraint. If you mix them before each one is finished, you will end up fixing a music problem by touching the voice, or vice versa, and you will chase your own tail for hours.
A useful habit: build a narration-only version of your video and watch it once with the music muted. If it holds attention on its own, the music is enhancement. If it falls apart, you are using music to hide a script problem, and no bed will save it.
Writing Narration Prompts That Sound Human
Give the model a delivery brief, not just a script
Text is not a performance. If you paste a paragraph and press generate, you get a default reading: neutral energy, even pacing, occasional odd emphasis. The fix is to write a short delivery brief before the script, describing the role, the mood, and the target listener.
A brief that works looks like this:
Voice: warm, mid-range, conversational. Speaker sounds like an experienced teacher explaining something they find genuinely interesting, not a broadcaster reading copy. Pace: moderate, with real pauses at paragraph breaks. Energy: rises slightly on the key insight of each section, settles at the end of each. Avoid upward inflection at the ends of declarative sentences.
That last line matters more than beginners expect. Many models default to a rising terminal contour, which reads as uncertainty or as a question. Explicitly asking for falling endings on statements is one of the cheapest quality upgrades available.
Control pronunciation and numbers
Pronunciation errors are the fastest way to break the illusion of a real narrator. Three categories cause most of them:
- Proper nouns. Brand names, place names, and invented terms. Write them phonetically in the script, or supply a pronunciation hint next to the word on first use.
- Numbers and units. Decide whether you want "1,200" read as "one thousand two hundred" or "twelve hundred." Spell it out exactly as you want it spoken. Abbreviations like "approx." and "vs." should be expanded.
- Homographs. Words like lead, read, bass, and live change sound based on meaning. If the model trips on one, rewrite the sentence so context forces the right reading.
Keep a running pronunciation sheet for any series with recurring names. Re-solving the same word for every episode is wasted effort.
Multi-speaker dialogue without a casting session
Conversational content — interviews, explainer duos, dramatized scenes — needs two or more distinct voices that stay consistent. The reliable approach is to define a speaker profile once per character and reuse it verbatim:
- A name, a one-line persona description, and a delivery brief.
- Format the script with the speaker name at the start of each line.
- Keep each character's traits identical across every scene. Changing a single adjective changes the output.
Once profiles are fixed, generate one character at a time rather than having the model produce the whole exchange in one pass. Separate tracks give you the ability to fix a single line later without regenerating the scene, and they let you overlap lines slightly for a more natural back-and-forth rhythm.
Retakes: change one variable at a time
When a line does not sound right, resist the urge to rewrite everything. Change one thing — the emphasis word, the pause placement, the pace adjective — and regenerate. Log what you changed. After a few passes you will know which levers actually move the output for your voice, and you will stop burning time on random rewrites.
Generating Background Music That Follows the Emotional Arc
Map the arc before you generate
Do not open a music tool before you can describe your video's emotional shape in five to seven beats. Write them down with approximate timings:
- 0:00–0:12 — curious, restrained, sparse.
- 0:12–0:45 — patient explanation, light rhythmic pulse entering.
- 0:45–1:20 — building tension around the problem.
- 1:20–1:55 — the turn: warmer, more open, higher register.
- 1:55–2:30 — confident, driving, full arrangement.
- 2:30–end — resolution, thinning out, no hard stop.
That map becomes your prompt plan. Generate one cue per beat rather than asking for a single track that morphs. Short cues edited together give you far more control than one long piece, and they are easier to replace individually.
Prompt the texture, not just the genre
"Cinematic music" produces generic results. Describe the instrumentation, density, and energy instead:
- Instrumentation: felt piano, muted pizzicato strings, single sustained synth pad, brushed drums, plucked harp.
- Density: solo instrument with room noise, or three to five elements maximum, or full ensemble.
- Motion: static drone, slow swell every eight bars, steady sixteenth-note pulse.
- Register and brightness: low and dark, mid-range and neutral, bright and airy.
- Mix notes: no lead melody over the voice, no percussion in the first twenty seconds, no vocal samples.
That last category is the one people forget. If your music has a melodic hook that competes with the narration, no amount of level adjustment will fully solve it. Ask for texture rather than tune when the track sits under speech.
Genre blending is another lever. "Lo-fi hip-hop tempo with orchestral strings instead of drums" or "ambient synth bed with a light acoustic guitar figure" gives you something specific enough to be interesting but generic enough to stay out of the way.
Stems, loops, and the eight-bar rule
If your tool can export stems — drums, bass, melodic elements separately — always take them. Stems let you drop the percussion for a quiet passage, keep the pad, and bring the drums back at the turn. That kind of dynamic arrangement is what makes a soundtrack feel composed rather than dropped in.
When you only have a stereo file, you can still fake this: generate the same cue twice with different density prompts and cross-fade between them at scene changes. The ear reads it as an arrangement change.
Aim to keep musical phrases in multiples of eight bars wherever you cut. Landing a scene change on the downbeat of a new phrase feels intentional; landing mid-phrase feels like a glitch.
Hybrid approach: library music plus generated beds
Generated music excels at matching an unusual brief and at being genuinely yours. Licensed library music excels at polished production value and predictable structure. Use both.
A practical split:
- Generated cues for transitions, stings, tension sections, and any moment that needs to line up exactly with a specific visual beat.
- Library tracks for long, low-attention stretches where a well-produced loop simply works.
Match the two by tonality. Write down the key of your main library track and ask for generated cues in a compatible key, or pitch-shift cues slightly to fit. Consistent tonal center across a video is the difference between a soundtrack and a playlist.
Sync and Mix: Making Audio and Picture Agree
Ducking and the speech pocket
Ducking — automatically lowering music when narration plays — is standard, and it is also commonly overdone. If you push the music down 12 dB every time a sentence starts, the bed will audibly pump.
Better settings for spoken-word video:
- Static reduction of 6–9 dB on the music bus, applied for the whole video, so you are never fighting the vocal.
- Additional dynamic ducking of only 3–4 dB under speech, with a slow attack (around 20–40 ms) and a slow release (300–600 ms).
- Sidechain the music to the voice track so it reacts only to the actual narration, not to room tone.
If you find yourself needing more than 12 dB of total reduction, the music is simply too busy. Change the cue.
Hit points and beat alignment
A hit point is any moment where audio and picture agree on purpose: a cut, a reveal, a text-on-screen, a camera move. Aligning at least a few of these per minute to a musical accent creates the impression that the whole piece was scored by hand.
Workflow: place markers on your timeline at the five or six most important visual moments. Then slide the music cue so its strong accent — usually the first beat of a phrase — lands on the most important marker. Nudge the remaining markers to approximate agreement. Perfection is not required; the brain is generous when one or two moments lock cleanly.
Room tone and transitions
Digital silence is a tell. When narration stops and the music has not started, a listener hears a hole. Fill those gaps with two to three seconds of quiet room tone or a very low ambient pad at -45 to -50 dB.
For transitions between musical cues, avoid butt-splicing. Overlap the outgoing and incoming cues by half a second to two seconds, and let the reverb tails cross. If the two cues are in different keys, put the transition on a percussive moment, or add a short riser to mask the change.
The Cleanup Pass: Loudness, EQ, and Intelligibility
EQ: carving space
Narration and music occupy overlapping frequency ranges. Rather than boosting the voice, carve the music:
- Voice: high-pass at 80–100 Hz to remove rumble. Gentle presence lift around 2–5 kHz only if the voice sounds dull. Narrow cut around 200–400 Hz if it sounds muddy or boxy.
- Music: high-pass at 100–150 Hz when under speech, and a broad 2–4 dB dip between 1 and 4 kHz — the intelligibility band. This is the single most effective move in spoken-word mixing.
Make these cuts on the music bus, not on each cue, so the relationship stays constant across the whole video.
Compression: consistency without pumping
Narration recorded across multiple generations will have level and tone variation between lines. A gentle compressor — ratio 2:1 to 3:1, threshold set so you are seeing 3–6 dB of gain reduction on louder phrases, attack 10–20 ms, release 100–200 ms — evens that out. Follow it with 2–3 dB of makeup gain.
If the voice starts to sound flat or the room noise swells between sentences, you have compressed too hard. Back off and fix the offending lines individually.
Loudness targets per platform
Loudness normalization is why your video can sound quieter than the one next to it even though both peak at the same level. Integrated loudness targets that work across most platforms:
- General web and social video: -14 LUFS integrated, true peak no higher than -1 dBTP.
- Podcast and audio-first content: -16 LUFS integrated, -1 dBTP.
- Broadcast-leaning deliverables: -23 LUFS integrated, -2 dBTP.
Measure the full mix, not just the voice. If your mix reads -11 LUFS, the platform will turn it down and your carefully balanced voice-to-music ratio will survive — but any clipping will not. Leave the headroom.
Noise, sibilance, and plosives
Three targeted fixes handle most defects:
- Sibilance (harsh s and sh sounds): a de-esser in the 5–8 kHz range, reducing 3–5 dB only on the sharpest peaks.
- Plosives (pops on p and b): a high-pass filter, or manual volume automation pulling 2–4 dB for 30 ms at the moment of the pop.
- Background hum: a narrow notch at 50 Hz or 60 Hz plus harmonics, depending on your region's mains frequency.
Do these before compression, or the compressor will amplify the problems you are trying to remove.
A Repeatable Eight-Step Workflow
- Script and mark up. Write the narration, insert pause and emphasis cues, expand numbers, and flag difficult proper nouns.
- Map the emotional arc. Five to seven beats with timings.
- Generate narration line by line. One speaker at a time, consistent voice profiles, retakes logged.
- Assemble the voice track. Cut and space lines, add room tone, listen once with no music.
- Generate music cues. One per beat, texture-first prompts, stems where available.
- Place and align. Set cue boundaries on phrase starts, align hit points to key visual moments.
- Mix. Music EQ dip, static reduction, sidechain ducking, transition overlaps.
- Master and verify. Compressor, loudness measurement, true-peak check, phone-speaker listen.
Steps four and eight are where most projects gain or lose their polish. Neither takes long, and both are routinely skipped.
Common Mistakes and How to Fix Them
Music too loud under speech. The most frequent error. If you can reliably transcribe the music's melody while narration plays, it is too loud or too busy. Cut the 1–4 kHz band before you cut overall level.
One long generated track for the whole video. It will not match the arc, and you cannot fix the mismatch without regenerating everything. Use short cues.
Inconsistent voice across scenes. Caused by rewording the voice brief between sessions. Save profiles and reuse them verbatim.
Hard music stops at the end. Music that ends on a hard cut feels unfinished. Generate a one-bar tail or fade over six to ten frames.
No room tone in gaps. Digital silence between narration blocks sounds broken. Fill with a low ambient layer.
Loudness measured on the voice only. Platforms normalize the whole mix. Measure the master.
Over-de-essing and over-compressing. Both make speech worse, not better. Apply the minimum that solves the problem.
Choosing the Right Tools
When evaluating AI voice and music tools for a recurring workflow, weight these criteria in order:
- Consistency. Can you save a voice profile and get identical results a month later? Voice drift is the most expensive problem to fix after the fact.
- Emotional control. Look for per-line direction — pace, emphasis, pause — not just a global "style" selector.
- Stem export. Separate music elements turn a rigid track into an adjustable arrangement.
- Pronunciation control. Custom lexicons or phonetic overrides save hours on technical and branded content.
- Licensing clarity. Confirm commercial use, redistribution rights, and whether outputs are exclusive to you. Keep the terms on file.
- Multilingual output. If you localize, the same voice across languages preserves brand identity.
- Export formats. WAV at 48 kHz for finishing; compressed formats only for review copies.
Test any candidate tool on a real 60-second script before committing. Demo reels hide pacing problems and pronunciation failures that show up immediately on your own material.
FAQ
How do I stop AI narration from sounding flat?
Add a delivery brief per section, not per video. Vary pace and pause length between paragraphs, and explicitly request falling intonation at the end of statements. Most flatness comes from a uniform reading with no dynamic variation.
Should I generate one music track or several?
Several. Short cues of 15 to 45 seconds each, generated per emotional beat, give you control that a single three-minute track never will.
Do I need to duck the music automatically?
Static reduction plus gentle sidechain ducking works best. Heavy automatic ducking produces audible pumping, which is more distracting than a slightly loud bed.
What loudness should I target?
-14 LUFS integrated with a -1 dBTP ceiling covers most web and social platforms. Podcasts typically sit at -16 LUFS.
Can I mix generated and licensed music in one video?
Yes, and it is often the best approach. Match tonality between them, and use generated cues at the moments that need precise sync.
How long should an AI voice cue be before I regenerate instead of edit?
If fixing a line takes more than two minutes of editing, regenerate it. Editing artifacts accumulate faster than most people expect.
Is AI narration acceptable for client work?
Increasingly, yes, provided the script is strong and the mix is finished properly. Disclose your production method when a contract requires it.
Ship the Audio Checklist
Before you export, run this list once. It takes four minutes and catches nearly everything.
- Narration is intelligible on a phone speaker at 50% volume.
- Music drops 3–4 dB under speech without audible pumping.
- Every scene transition lands on or near a musical phrase boundary.
- No gap between voice and music contains digital silence.
- Music fades or resolves rather than cutting off.
- Integrated loudness is at target and true peak is below the ceiling.
- Sibilance and plosives are controlled by targeted fixes, not global compression.
- The voice sounds like the same person from the first line to the last.
Audio is the cheapest place in a production to look professional. It does not need a studio, a composer, or a voice actor to get right — it needs a defined arc, a delivery brief, an EQ dip in the intelligibility band, and a loudness check at the end. Get those four habits into your workflow and the synthesized elements stop sounding like shortcuts and start sounding like the reason your videos hold attention.



