Why Audio Quality Decides Whether Viewers Stay
Audiences forgive imperfect visuals far more readily than they forgive bad sound. A slightly soft focus, a shaky handheld shot, even an overexposed window in the background — viewers will tolerate all of it if they can hear clearly and the tone feels intentional. But the moment narration sounds robotic in the wrong way, music fights the voice, or loudness jumps between scenes, the video starts to feel amateur. On mobile, where most short-form viewing happens, that perception forms in the first three seconds.
This is why audio deserves its own production pass rather than being bolted on at the end. A video editing timeline that is already locked to the frame is a poor place to discover that your narration reads sixty seconds too long, or that the music bed you generated has no clean loop point for the final scene. Treating narration, music, and effects as three coordinated layers from the start saves hours of re-editing later.
This guide lays out a complete, tool-agnostic workflow for producing AI narration and AI-generated background music, then mixing them so the result sounds deliberate. Whether you publish tutorials, product demos, documentaries, social clips, or course material, the same principles apply: write for the ear, cast the voice deliberately, generate music that serves the scene, and mix to known loudness targets.
The Three Layers of a Video Soundtrack
Before touching any tool, separate the soundtrack into three functional layers. Each has a different job, and confusing those jobs is the root of most bad mixes.
Narration: the information layer
Narration carries meaning. It should be intelligible at low volume, on a phone speaker, in a noisy room. Every decision about it — pacing, diction, sentence length — serves comprehension first and aesthetics second. If a listener has to concentrate to understand the words, nothing else in the mix matters.
Music: the emotion layer
Music tells the viewer how to feel about what they are seeing and hearing. It also hides small imperfections in editing rhythm by smoothing transitions. Crucially, music is not supposed to carry information. If the viewer notices the music, it is either doing something wonderful or something wrong.
Effects and ambience: the realism layer
Room tone, wind, keyboard clicks, whooshes on cuts, subtle risers before reveals. This layer is optional but disproportionately powerful. A narration track with zero ambience sounds sterile and obviously synthetic; adding a light room tone underneath instantly makes it feel like a real recording.
A useful rule: mix in order of importance — narration first, then music, then effects. Never set music levels before the voice is sitting where you want it.
Writing Narration Scripts That AI Voices Read Well
Text-to-speech has improved dramatically, but it still reads what you give it literally. Scripts written for the eye often fail in the ear. A few adjustments close most of the gap.
Punctuation is your direction track
Commas create micro-pauses. Periods create full stops. Em dashes create hesitation. Colons create a small lift before the payoff. If a sentence feels rushed when you preview it, add a comma at the natural breath point rather than rewriting the sentence. If it feels sluggish, remove one.
Paragraph breaks matter too. A blank line between two thoughts usually produces a slightly longer pause than a period, which is useful for section transitions without inserting awkward silences manually.
Sentence length and rhythm
Alternate long and short sentences. Long sentences build context; short ones land the point. A script made entirely of medium-length sentences — the default of business writing — produces a hypnotic monotone, regardless of how expressive the voice model is.
A practical test: read the script aloud and tap the table on each stressed syllable. If your taps fall into a perfectly even metronome, rewrite.
Numbers, names, and jargon
Numbers are the single most common source of weird AI narration. "2024" might be read as a year or as a quantity depending on context. Product names with unusual casing get mangled. Acronyms get spelled out letter by letter when they should be spoken as words, or the reverse.
Three fixes work reliably:
- Write numbers the way you want them spoken ("two thousand twenty-four" or "twenty twenty-four").
- Add phonetic spellings for names the model keeps missing ("Ka-REEN" instead of "Karine").
- Expand acronyms on first use if the spoken form should be letters ("S-E-O," not "seo").
Length budgeting
As a planning figure, expect roughly 140 to 160 spoken words per minute at a comfortable documentary pace, and 170 to 190 for energetic explainer content. A five-minute video therefore needs about 750 to 900 words of narration. Writing to that number before you record prevents the classic problem of cutting content after the visuals are locked.
Choosing a Voice: A Practical Decision Framework
Voice selection is a casting decision, and casting decisions are best made against criteria rather than vibes.
What to evaluate in a sample
Generate the same 80 to 120 words of your actual script with three or four candidate voices, then compare:
- Intelligibility at low volume. Play it at 30 percent on a phone speaker. Can you still follow it?
- Consistency across the whole sample. Does the pitch drift or the energy sag halfway through?
- Emotional range. Does it handle a question, a warning, and a light joke without sounding identical in all three?
- Pacing control. Can you slow it down 10 percent without artifacts?
- Plosive and sibilant handling. Hard P and B sounds should not thump; S sounds should not hiss.
Matching voice to format
| Content type | Voice profile that works | Voice profile that fails |
|---|---|---|
| Product demo | Neutral, confident, moderate pace | Overly warm or theatrical |
| Documentary | Slower, deeper, subtle dynamics | Bright, fast, salesy |
| Social short | Energetic, higher pitch, punchy | Slow, formal, monotone |
| Course module | Clear, even, slightly slower | Dramatic, inconsistent energy |
Consent, cloning, and disclosure
If you clone a voice, use your own or one you have explicit written permission to reproduce. For voices of public figures, avoid imitation entirely. Many platforms and regulators now expect disclosure when synthetic narration represents a real person, and audience trust erodes fast when imitation is discovered. A one-line on-screen note or a line in the description is cheap insurance.
Prompting Background Music That Actually Matches the Scene
Music generation models respond best to descriptions of function and texture, not just genre labels.
Describe the job the music must do
Instead of "cinematic orchestral," try "calm, hopeful underscore for a slow product reveal, no percussion, sparse piano with warm strings, leaves space for narration." The second prompt contains the information the model actually needs: energy level, instrumentation, density, and a constraint about the mix.
The four dials worth specifying
- Tempo. Slow (60–80 BPM) for reflection, mid (90–110) for momentum, fast (120+) for energy. State it explicitly.
- Density. How many instruments are playing at once. Low density leaves room for voice.
- Register. Low, mid, or high frequency focus. Mid-heavy music collides with the human voice; low and high sit better underneath it.
- Arc. Does the track build, hold steady, or resolve? A steady bed is usually easier to edit than a dramatic arc you have to cut around.
Loops, stems, and edit headroom
Ask for loopable or seamless results whenever possible, and prefer generated tracks with stems if your tool offers them. Stems let you mute the melody during narration and bring it back in the gaps — a technique that makes a single music bed feel custom-scored.
If stems are not available, generate a track that is 15 to 20 seconds longer than you need and edit inside the sustained sections rather than at the climactic moments. Cutting on a downbeat is easier than cutting through a swell.
Mixing: Levels, Ducking, and Loudness
The basic balance
Start with narration peaking around −6 dBFS and averaging around −16 to −18 dBFS. Place the music bed so that it sits roughly 15 to 20 dB below the narration during spoken passages. That is quieter than most beginners expect; it is also what makes words effortless to follow.
Ambience and effects sit between the two, usually 20 to 25 dB below narration, loud enough to be felt rather than heard.
Sidechain ducking, step by step
Ducking automatically lowers music when narration plays. If your editor supports sidechain compression, the setup is quick:
- Put narration on its own track and music on another.
- Add a compressor to the music track and set its sidechain input to the narration track.
- Set threshold so ducking triggers only on speech, not on room tone.
- Set ratio around 4:1 to 8:1.
- Aim for 6 to 10 dB of gain reduction.
- Set attack around 5–10 ms and release around 200–400 ms so the music breathes back naturally.
If sidechain is unavailable, draw volume automation manually instead. Manual automation gives finer control and is often the better choice for a two-minute video with only a dozen speech segments.
Loudness targets
Deliverable loudness varies by destination:
- Online video platforms: roughly −14 LUFS integrated, true peak no higher than −1 dBTP.
- Podcast and audio-first distribution: −16 LUFS.
- Broadcast television in Europe: −23 LUFS; in North America, −24 LKFS.
Measure integrated loudness on the finished mix, not on individual tracks, and check true peak after any limiter. If your mix hits −7 dBTP on a phone, listeners will hear crackle and will blame your content, not their device.
A Repeatable End-to-End Workflow
Here is the sequence that keeps audio from becoming the bottleneck on every project.
- Write the script for the ear. Read it aloud, adjust punctuation, and confirm the runtime against your target length.
- Audition voices on real script lines. Pick two finalists, then decide with headphones and with a phone speaker.
- Generate narration in sections. Short paragraphs give you edit points and make re-recording a single line cheap.
- Clean the voice. Light noise reduction, high-pass filter around 80–100 Hz, gentle de-essing if needed, and a subtle compressor for evenness.
- Generate music per scene, not per project. Three or four short beds matched to emotional beats edit better than one long track.
- Lay ambience. Room tone under narration, plus scene-specific effects for cuts and transitions.
- Duck and balance. Set narration first, duck the music under it, then bring effects in.
- Master to target loudness and export at a sensible bitrate — 192 kbps AAC or higher for stereo delivery.
The order matters more than any single tool. Skipping step one guarantees rework in step three; skipping step seven guarantees a mix that sounds fine in headphones and muddy on a phone.
Common Mistakes and How to Fix Them
Narration sounds flat and lifeless. Usually a script problem, not a voice problem. Add sentence variety, split long clauses, and regenerate only the flat sections rather than the whole file.
Music competes with the voice. Check the frequency range first — if the bed is busy between 300 Hz and 3 kHz, high-pass it or choose a sparser arrangement. Then check level: if ducking is under 4 dB, it may as well not be there.
Volume jumps between scenes. Almost always caused by generating each narration segment at different settings. Normalize all voice segments to the same target before assembling.
Sibilance or harshness on S sounds. Apply a de-esser before any broadband compression, not after. Compression amplifies the problem otherwise.
The ending feels abrupt. Add a two- to four-second musical tail and let the ambience fade under it. Silence after a hard cut reads as an error.
Everything sounds synthetic. Add imperfection deliberately: a slight room tone, one breath before a key sentence, a small timing offset for effects. Perfectly clean audio is the tell.
Where Each Type of Tool Fits
A practical stack for AI-assisted audio usually includes four roles, and one tool can sometimes cover two of them:
- Script and pacing tools for measuring runtime and flagging tongue-twisters before generation.
- Speech synthesis for narration, preferably with pitch and pace controls and per-line regeneration.
- Music generation for scene-matched beds, ideally with tempo and instrumentation controls.
- A digital audio workstation for cleanup, ducking, loudness metering, and export.
If you are working inside an all-in-one video editor, check whether it supports sidechain compression and LUFS metering. Many consumer editors do not, in which case export the audio, finish it in a dedicated tool, and reimport the mixed track.
For teams producing volume content, standardize on a template: fixed track names, a saved ducking preset, and a loudness target baked into the export settings. Templates turn a creative task into a repeatable one, which is the only way audio quality survives a busy publishing schedule.
Pre-Publish Quality Checklist
Run this before every upload:
- Narration is intelligible on a phone speaker at 30 percent volume.
- No word is clipped at the start or end of a segment.
- Music sits 15–20 dB below narration under speech.
- Ducking is smooth, with no audible pumping.
- Integrated loudness is within 1 LU of your platform target.
- True peak is at or below −1 dBTP.
- No unintended silence longer than half a second except where designed.
- Ambience continues under scene transitions rather than cutting abruptly.
- The final four seconds resolve musically rather than stopping dead.
- Any synthetic voice disclosure required by your platform or jurisdiction is present.
FAQ
How long should my music bed be?
Generate 15 to 20 seconds longer than the final runtime so you can trim from the middle and still have clean entry and exit points.
Should narration or music be generated first?
Narration. Its rhythm, length, and emotional tone determine what the music must do. Generating music first often produces a beautiful bed with nowhere to sit.
Can I use one voice model for an entire series?
Yes, and consistency is usually an advantage for brand recognition. Keep a saved preset with pitch, pace, and style settings documented so every episode matches.
How do I make AI narration sound less robotic?
Three levers, in order of impact: better script punctuation, shorter generation segments, and light post-processing with compression and a touch of room tone. Rewriting the script beats any plugin.
What loudness should I target if I publish to multiple platforms?
Master to the quietest relevant target, usually −16 LUFS, then let each platform's normalizer raise it. Mastering loud and hoping platforms turn it down is riskier than the reverse.
Do I need stems, or is a finished stereo track enough?
For simple talking-head content, stereo is fine. For anything with dramatic scene changes, stems or per-scene beds give you far more control with much less automation fiddling.
How much time should audio take relative to editing?
For explainer-style content, plan on roughly one-third of total production time for script, narration, music, and mixing. Teams that budget less end up spending it anyway — during re-edits.


