Why audio decides whether a viewer keeps watching
Most of the effort in an AI video project goes into the picture: prompt tuning, camera motion, color, pacing of cuts. Audio gets whatever time is left. Viewers behave in the opposite way. They will forgive a slightly soft render, a strange hand, or a background that shifts unnaturally between shots. They will not forgive a narrator that has to shout over a music bed, or a soundtrack that swells at the exact moment the most important sentence lands.
There is a simple test worth internalizing: play your video from another room with the screen off. If the meaning still arrives, the mix works. If you can hear music but not words, the mix has failed no matter how good the visuals are. That single test reframes audio as the load-bearing element of comprehension rather than decoration added at the end.
The practical problem is that an AI-narrated video is always a two-source mix, and usually a three-source one. A synthetic voice has no room, no breath noise, and no natural dynamics, so it does not glue itself to music the way a recorded human voice does. Music generated or selected from a library has no idea what the script is saying. Ambience is often missing entirely. The job is not to turn knobs until it sounds nice; it is to build a small system that produces a consistent result every time you publish.
The three-layer model: voice, music, ambience
Treat every AI video as three layers with clearly defined responsibilities. Ambiguity between layers is where most amateur-sounding results come from.
What each layer is responsible for
- Voice: carries the meaning. Everything else must yield to it. Target: always intelligible on a phone speaker at 40% volume.
- Music: carries emotion and momentum. It should reinforce what the script already implies, never introduce a competing idea.
- Ambience and room tone: carries continuity. A quiet bed of air, room, or texture prevents the mix from feeling like two unrelated files stacked on top of each other, and it smooths transitions between chapters.
A useful rule of thumb: if you mute the voice and the music alone tells a different story than the script, the music is wrong. If you mute the music and the video feels emotionally flat, the music is doing its job.
Set loudness targets before you mix
Decide numbers first, then mix toward them. Streaming platforms normalize playback to roughly -14 LUFS integrated, podcasts often sit near -16 LUFS, and true peak should stay at or below -1 dBTP to survive lossy encoding. Under narration, music usually sits 18 to 24 dB below the voice, with momentary ducking taking it down another 4 to 7 dB.
Peak meters tell you almost nothing useful here. Use an integrated loudness meter on the master bus and a short-term meter while you ride levels. Metering first, taste second, is what separates a repeatable workflow from a session of guessing.
Step 1: Analyze the voice track before touching the music
You cannot match music to a voice you have not listened to carefully. Play the narration once with your eyes closed and take notes.
Read tone, pace, and emphasis
Write down four things: pace in words per minute, pitch range (narrow or wide), emotional register (calm authority, friendly explainer, urgent promo, dry documentary), and where the script pauses. A calm explainer at roughly 150 to 160 words per minute with a narrow pitch range pairs well with ambient pads, soft mallets, and low percussion. A promo at 180 to 190 words per minute with wide dynamics and hard consonants needs a percussive bed with sparse melody and no vocal samples.
Also note articulation problems: swallowed word endings, sibilant hiss, clicks, and unnatural emphasis placed on the wrong syllable. These are script and generation problems, not mixing problems, and no EQ will fix them.
Clean and level the voice first
Fix the voice before any music enters the session:
- Trim silences to 0.25 to 0.4 seconds. Longer gaps kill pacing; shorter gaps make the narration feel breathless.
- High-pass at 80 to 100 Hz to remove rumble that will only fight the bass of the music.
- De-ess between 5 and 9 kHz with a narrow dynamic cut rather than a static EQ notch.
- Apply gentle compression, roughly 3:1 with 3 to 6 dB of gain reduction on louder passages.
- Normalize peaks to about -3 dBFS and leave headroom.
If a sentence sounds robotic, split it into shorter lines and regenerate that line only, then rejoin. Matching room tone across regenerated lines matters more than perfect timbre matching; a 5 ms crossfade at the edit point usually hides the seam.
Step 2: Choose music that matches energy, not just genre
Searching a music library by mood tag alone produces generic results. Work from an energy map instead.
Map the energy curve of the script
Make a simple beat sheet before you search for anything:
| Timecode | Narrative beat | Energy (1-5) | Music role |
|---|---|---|---|
| 0:00-0:08 | Hook question | 3 | Subtle pulse under first line |
| 0:08-0:40 | Problem framing | 2 | Pad, no drums |
| 0:40-1:20 | Solution walkthrough | 3 | Add soft percussion |
| 1:20-1:30 | Key proof point | 4 | Full texture, melody enters |
| 1:30-2:00 | Recap and close | 2 | Strip back, resolve on a chord |
Now choose tracks whose own energy curve roughly mirrors this shape. A great track with the wrong shape will force you to fight it with volume automation for the entire edit, which never sounds as good as choosing correctly the first time.
Tempo, key, and instrumentation
Tempo should support the voice rather than compete with it. As a starting point, keep music between 90 and 110 BPM under dense narration, and 120 to 140 BPM under sections where the viewer is mostly reading text on screen. Listen for whether the music's rhythmic accents land between your narrator's syllables or on top of them. Accents that collide with stressed words create a constant, low-grade irritation that viewers feel but cannot name.
Avoid tracks with prominent lead vocals under narration, and avoid strong melodic motifs during the sections where you need the viewer to remember a specific fact. Instrumentation choices do a lot of the work: pads, mallets, and plucks under explanation; drums, brass, and staccato strings for reveals; solo piano or low strings for emotional resolution.
Prefer stems and loopable edits
When the option exists, work with stems separated into drums, bass, melody, and pads. This lets you drop the melody for eight seconds during a critical claim without losing the rhythmic bed. Look for tracks delivered as clean loops without baked-in fades, so you can extend or shorten sections to match the script rather than cutting the script to match the music.
Step 3: Ducking, EQ carving, and gain staging
The mix itself comes down to three moves: carve space in the music, duck it when the voice speaks, and keep gain structure clean.
Sidechain ducking that does not pump
Set up a voice bus and use it to key a compressor on the music bus. Reasonable starting values: 4 to 7 dB of gain reduction, attack 10 to 20 ms, release 250 to 500 ms, and a lookahead of a few milliseconds if your tool offers it. The failure mode is a ducked bed that audibly breathes in and out between every sentence. If you hear pumping, lengthen the release so the music recovers slowly through short pauses instead of pumping back up.
For dense narration where the music never really breathes, dynamic EQ works better than full-band ducking. Duck only the 1.5 to 4 kHz region by 3 to 4 dB when the voice is present. The low end and air stay intact, so the bed still feels full while speech remains clear.
Carve a mid-range pocket
A synthetic voice usually has a fundamental around 100 to 200 Hz and intelligibility concentrated between 1.5 and 4 kHz. Carve accordingly:
- Broad 2 to 4 dB dip in the music between 1.5 and 4 kHz, either static or dynamic.
- Gentle high-pass on the music at 60 to 80 Hz so the voice's chest tone is not competing with a wall of sub energy.
- Small 1 to 2 dB dip around 200 to 300 Hz on the music to prevent mud, especially with pad-heavy tracks.
- If the voice sounds thin, boost 2 dB near 150 Hz rather than turning up the whole track.
Gain staging and bus structure
Build a simple hierarchy: voice, music, ambience, and effects buses all feeding a mix bus, then a master. Typical chain per bus is high-pass, corrective EQ, compressor, then a limiter on the master only. Keep the mix bus peaking near -6 dB before the master limiter, and keep the master limiter doing less than 2 dB of gain reduction. A master limiter working hard is a sign that you pushed levels instead of building them.
Step 4: Transitions, stings, and pacing
Transitions are where mixes either feel designed or assembled. Handle them deliberately.
Openings and chapter changes
For an explainer, fade the music in over 1.5 to 3 seconds starting slightly before the first word. For a hook-driven promo, hit the music on the first word with a short impact, and let the voice land on top of the decay. At chapter changes, use one of three moves: a two-second riser into the new section, a half-second drop-out just before the cut, or a hard texture change. Pick one per video and repeat it, so the viewer learns your structure without noticing it.
Use silence deliberately
A 0.6 to 1.2 second gap of full silence before a reveal is one of the most effective tools available, and it costs nothing. Equally useful are music-only bridges of two to four seconds that break up long narration stretches. The one thing to avoid is an abrupt music stop mid-note; end on a resolved chord or a 1.5-second fade instead.
Ride levels across long videos
For anything longer than two minutes, place music regions per chapter rather than running one track underneath from start to finish. Draw volume automation to hit your ducking targets, and keep an ambience layer running continuously so the room never appears to change between chapters. That continuous layer is what makes separately generated narration lines sound like one performance.
Keeping audio consistent across a multi-tool pipeline
Most AI video workflows touch several tools: one for scripting, one for narration, one for imagery, one for music, one for editing. Each handoff is a chance for the sound to drift.
Write a one-page audio style guide
Keep a living document with the values you have settled on: voice model and preset, target pace range, preferred music families, ducking preset, loudness targets, transition vocabulary, and a file-naming convention. Reusing the same voice and music family across a series is a bigger quality win than chasing the theoretically perfect voice for each episode. Consistency reads as professionalism.
Handle voice changes gracefully
If you must change narration voice mid-series, make the switch at a chapter boundary and crossfade ambience across the seam. Match pace first with a small time adjustment of 3 to 5 percent, then address timbre. Re-record the opening lines so the first 20 seconds establish the new voice clearly rather than sneaking it in.
Quality control: the listening pass
Before exporting, run the same checks in the same order every time.
The device matrix
Listen on a phone speaker, laptop speakers, earbuds, over-ear headphones, and if possible a car stereo. The phone speaker is the decisive test, since a large share of viewers watch that way. If words disappear on a phone speaker, the music is too loud in the 1.5 to 4 kHz range, regardless of what the meter says.
Common failure modes and fixes
- Pumping music: increase ducking release time, reduce ducking depth.
- Muddy dialogue: high-pass music at 60 to 80 Hz and dip 200 to 300 Hz.
- Harsh sibilance: dynamic de-esser on the voice, not a global high-shelf cut.
- Voice sounds distant: reduce the music's reverb or add a short 20 to 40 ms slap delay to the voice for presence.
- Music too loud at the top: automate the first five seconds 3 dB lower than the rest.
- Inconsistent loudness between chapters: measure each chapter, then match by ear, then re-measure the whole video.
- Distortion after export: check true peak before encoding, and lower master output by 1 dB.
Delivery specs and export habits
Export at 48 kHz stereo, AAC at 192 to 320 kbps, integrated loudness near -14 LUFS for video platforms, and true peak no higher than -1 dBTP. Keep a 24-bit WAV master and stem exports archived alongside the project, plus music-only and voice-only versions. Those extra versions cost almost nothing at export time and save hours when you need subtitles, localization, or a version of the video for a client who wants different music.
Resist the urge to make the mix louder to feel more professional. Platforms normalize playback, so extra loudness only reduces dynamic range. A mix that is 2 dB quieter with clear separation will always feel better than one that competes with itself.
FAQ
Should I write the script with music in mind?
Yes. Marking the two or three moments that need musical emphasis while writing means your beat sheet is already done by the time you start mixing. It also prevents the common mistake of scoring every sentence equally.
Is generated music better than a library track?
Generated music is easier to shape to an exact length and mood, but libraries tend to sound more coherent over long stretches. A practical hybrid is to use library tracks for the main bed and generated stingers for transitions.
How loud should background music be under narration?
Start 20 dB below the voice, then adjust by ear on a phone speaker. Sections with no dialogue can sit 6 to 10 dB below the voice instead, since nothing is competing for attention.
What is the single highest-impact fix?
Put an ambience layer under the entire video. It costs one track, almost no processing, and it removes the sterile, stitched-together quality that synthetic narration tends to have.
How do I keep a series sounding consistent?
Pick one voice preset, one music family, and one ducking preset, and document them. Changing tools is fine; changing sonic identity between episodes is what breaks the audience's sense of place.


