Why audio quality decides whether a video feels professional
Audiences forgive a lot. They forgive slightly soft focus, a handheld frame that drifts, a background that is not perfectly dressed. What they rarely forgive is audio that sounds thin, uneven, or out of sync. Bad audio reads as amateur within two seconds, and it also costs comprehension: when a viewer has to work to understand a voiceover, they stop watching. That single fact is why a repeatable narration and music pipeline matters more than chasing whichever tool is trending this month.
Generative voice tooling has changed the economics of narration. A small team can now produce a narrated explainer, a localized product tour, or a documentary-style short without booking studio time for every script revision. Music generation has followed the same path — describe a mood, get a usable bed in seconds instead of hunting through a licensing catalog.
But generation is the easy part. The hard part is making generated audio sit inside a finished edit without sounding synthetic, fighting the voice, or blowing out on one platform while vanishing on another. This guide walks through the whole pipeline: planning, script preparation, voice selection, music sourcing, mixing, localization, and quality control. Treat it as a workflow you can repeat rather than a one-off experiment.
Start with an audio map, not with the tools
Before you open a single generator, write down the audio decisions that are expensive to change later. Most rework in video projects comes from discovering, halfway through the edit, that the voice you locked does not support the runtime, or that the music bed you love cannot be licensed for the client's channel.
| Decision | Lock it early by | Why it matters later |
|---|---|---|
| Narration language | Confirming the primary market and any secondary ones | Changing language after the edit means re-timing every cut |
| Number of voice roles | Reading the script once, out loud, and marking speaker changes | Voice consistency is easier than matching a second voice later |
| Tone and register | Comparing two extremes: warm-conversational vs. authoritative | Tone drives pacing, music choice, and even the edit rhythm |
| Delivery platforms | Listing every channel the video will appear on | Loudness targets differ between social, streaming, and broadcast |
| Runtime budget | Estimating words per minute at your chosen pace | Script length is the root cause of rushed or padded audio |
A useful rule: narration at a comfortable, clear pace lands between 130 and 155 words per minute. Technical content with numbers and product names sits closer to 130. Listicles and social spots can push 165. If you write 1,400 words and you have a 90-second slot, no voice engine on earth will save that script.
Once the map is on paper, the rest of the work becomes execution rather than negotiation.
Script preparation: the highest-leverage step
Write for the ear, not for the eye
Readers skim; listeners cannot. That means:
- One idea per sentence. If a sentence has three clauses and a parenthetical, split it.
- Subject first, action second. "The panel opens on the left" beats "On the left is where the panel will be found."
- Concrete nouns over abstractions. "A 12-second clip" is easier to hear than "a brief asset."
- Numbers spelled the way you want them spoken. If the voice should say "twenty-five," write twenty-five. If it should say "two five," write 2-5.
Mark up the script for the voice engine
Synthetic voices respond well to structure. A script that looks like a wall of prose produces flat delivery because the engine has no cues about intent. Add lightweight markup:
- Paragraph breaks at every topic shift, not every two sentences.
- Commas where you want a micro-pause, periods where you want a full stop.
- Ellipses for hesitation or suspense, used sparingly.
- Phonetic spellings for brand names the engine will mispronounce on the first pass.
- A short pronunciation note at the top for acronyms, region names, and product SKUs.
Keep a project glossary. The fifth time you correct the same brand name, you will wish you had one.
Build the script in beats
Split the script into beats of 15 to 30 seconds. Beats make review manageable, and they map cleanly onto the timeline: one beat per scene, one beat per visual sequence. When a client asks for a rewrite, you regenerate one beat instead of the whole voice track. That is a tenfold productivity difference on revisions alone.
Selecting a synthetic voice that fits the job
Voice selection is not about which voice sounds most impressive in a demo. It is about which voice survives your specific script, at your runtime, on your audience's speakers. Demo reels are usually engineered with flattering text.
Build a shortlist with one shared test paragraph
Take a 90-second excerpt of your real script — ideally one that includes numbers, a proper noun, a question, and an emotional turn — and run the same excerpt through every candidate voice. Listen on three setups: headphones, a laptop speaker, and a phone speaker. Phone speakers are unforgiving and they are how most social viewers will actually hear you.
Score each candidate on:
- Clarity at 1x speed. Can you follow every word without the transcript?
- Natural breath. Does it pause where a person would breathe, or does it run out of air mid-clause?
- Emotional range. Can it shift from explanatory to warm without sounding like a different person?
- Pronunciation control. How easy is it to fix a mispronounced term?
- Consistency across takes. Regenerate the same line five times. Do the results drift?
That last test matters more than most people expect. A voice that varies between takes will fight you during revisions.
Match voice to format, not to preference
A voice that works for a 12-minute tutorial often collapses in a 15-second vertical ad, where every syllable has to land. Conversely, high-energy ad voices feel exhausting over a long course module. Choose based on the format you produce most, then build a second voice as an alternate for a different format — not a third and fourth. Audiences remember a voice, and rotating too widely erodes the identity you are building.
Direct the performance with pacing, emphasis, and pauses
Getting a good read is mostly about direction:
- Pacing. Generate 5 to 10 percent slower than your target, then tighten on the timeline if needed. It is far easier to remove silence than to add warmth.
- Emphasis. If the engine supports emphasis tags, use them on one word per sentence at most. Emphasizing everything emphasizes nothing.
- Pauses. Insert explicit silence rather than relying on punctuation. A 400 to 700 millisecond gap before a key reveal does more for retention than any musical swell.
- Volume dynamics. Ask for a slightly lower level on asides and parentheticals. That variation is what makes a read sound directed rather than generated.
Save your settings as a preset. When a new episode or video arrives, you start from a known-good baseline instead of re-discovering it.
Background music: sourcing, structure, and practical hygiene
Music has three jobs in a video: set emotional context, smooth transitions, and cover small audio imperfections. It should never compete with narration.
Match the energy curve, not the genre
Do not pick a track because you like the genre. Map the video's energy curve first — where it opens, where tension builds, where the payoff lands, where it resolves — then look for a track whose own arc roughly matches. A track that builds beautifully but peaks at 40 seconds will fight a video whose climax is at 90 seconds.
Decide between a bed and a feature
- Bed: continuous, low-level, no melody that competes with speech. Ideal for tutorials, product walkthroughs, documentaries.
- Feature: melodic, foregrounded, used at transitions or as a standalone intro sting. Ideal for brand spots and social hooks.
- Silence: underused. Dropping music entirely for 5 to 10 seconds before a major point makes the point land harder than any swell.
Keep the paperwork clean
Whatever source you use, record for each track: source, license type, permitted platforms, whether monetized channels are covered, and expiry if any. Store it in a spreadsheet alongside the project file. A music folder without a licensing log is a liability, and the question almost always arrives months later when nobody remembers where the file came from. Generated tracks still deserve a line in that log, including the prompt and the tool version you used.
Mixing: levels, ducking, and the numbers that matter
Set dialogue as the anchor
Mix around the voice, never around the music. Start by normalizing narration to roughly -16 to -14 LUFS integrated for web delivery, then bring everything else up to it. If the voice is too quiet, lower the music — do not raise the voice into clipping.
Common integrated loudness targets:
| Delivery target | Integrated loudness | True peak ceiling |
|---|---|---|
| Social and web video | -16 to -14 LUFS | -1 dBTP |
| Streaming video platforms | -14 LUFS | -1 dBTP |
| Podcast and audio-first | -19 to -16 LUFS | -1 dBTP |
| Broadcast (EBU R128) | -23 LUFS | -1 dBTP |
Two decibels of variation between platforms is normal and acceptable. Ten is not.
Duck the music properly
The blunt approach — drop the music 20 dB the moment anyone speaks — sounds unnatural. A better duck:
- Reduce the bed by 4 to 6 dB under narration, not more.
- Use a slow attack of 100 to 200 ms so the duck is felt rather than heard.
- Use a release of 400 to 800 ms so the music breathes back up between sentences.
- Sidechain the music to a copy of the voice track rather than hand-drawing volume curves.
Carve out the frequency space
Music beds and voices overlap heavily between 200 Hz and 4 kHz. A gentle high-pass on the voice around 80 to 100 Hz removes rumble. A small dip of 2 to 3 dB in the bed around 1 to 3 kHz, plus a subtle notch where the voice is loudest, buys intelligibility without making the track sound hollow.
Finish with a limiter set to a -1 dBTP ceiling and check the true peak readings after export. Loudness normalization on the platform happens after your export, so anything you clip stays clipped.
Localization and multi-language versions
Generating a second language is fast; making it feel native is not. Budget time for these steps.
- Translate for speech, not for subtitles. Subtitle translations are compressed for reading speed. Spoken translations should be expanded for breath and clarity.
- Re-time, do not re-speed. Every language lands differently: Spanish and German often run 15 to 25 percent longer than English. Plan for shots that can flex, and regenerate instead of squeezing the audio faster.
- Keep the voice family. If possible, use voices that share a timbre across languages so the brand identity persists when a viewer switches languages.
- Localize on-screen text and units. A voice saying "twenty kilometers" with an on-screen "12 miles" looks careless, even if both are technically correct.
- Check culture-specific references. Idioms, holidays, and humor rarely survive a literal translation and tend to sound strange when synthesized.
A practical shortcut: lock the picture edit first, then produce all language versions from the same timeline so the timing issues are identical and fixable in one pass.
Quality control checklist before export
Run the same checklist on every project. It takes four minutes and prevents nearly all viewer complaints.
- Listen start to finish on headphones with your eyes closed. You will hear problems you cannot see on a waveform.
- Listen again on a phone speaker at low volume. If the voice is unintelligible there, it is too quiet relative to the music.
- Confirm no clipped samples: check true peak, not just sample peak.
- Verify that every pronunciation in your glossary is correct.
- Check that music starts and ends cleanly — no abrupt cuts, no audible loop seams.
- Confirm sync drift: watch the last 20 seconds of the video to catch cumulative offset.
- Confirm the file exports at the correct loudness target for the destination platform.
- Archive the project audio folder, the script version, and the licensing log together.
Common mistakes and how to fix them
Music too loud under speech. The most frequent complaint in every review. Drop the bed 3 dB, then 3 dB more, and see if you miss it. You usually will not.
One flat read for the whole video. Vary pacing between sections. Even a small tempo shift between the intro and the body keeps attention.
Regenerating the entire voice track for one word. Work in beats. Regenerating one beat and stitching it back in preserves consistency and takes a fraction of the time.
Ignoring room tone and silence. Constant sound is fatiguing. Drop the music for a beat before a key line, and let the pause do the work.
Skipping the phone-speaker test. Desktop editing environments are far more forgiving than a phone in a noisy room. The phone is the real-world standard.
Mixing before the script is final. Every script change invalidates your timing. Lock the words before you invest in the mix.
Forgetting alternate versions. Social cuts, silent-autoplay versions, and audio-only versions usually need different mixes. Plan for them at export time rather than re-opening the project a month later.
FAQ
How many words per minute should I aim for in narration?
Between 130 and 155 words per minute for most explanatory content. Technical scripts with numbers sit near 130; fast social spots can reach 165. Measure one of your finished reads to calibrate your own pace.
Can I mix a generated voice with a human narration in the same video?
Yes, but match the room and the level. Human narration usually needs a touch of compression and de-essing to sit beside a synthetic voice. Test a 20-second crossover before committing to the approach.
What loudness target should I export at?
For social and web, -16 to -14 LUFS integrated with a -1 dBTP ceiling covers nearly everything. Match the platform's published target when you know it, and keep variation under two decibels across a series.
How do I keep a series sounding consistent across episodes?
Save voice settings as presets, keep a pronunciation glossary, and reuse the same music palettes and mixing template. Consistency comes from templates, not from talent.
Is generated music good enough for client work?
Often, yes — for beds and transitions. For a hero brand moment where music carries the emotion, a composed or carefully licensed track still tends to hold up better. Judge each case on whether the music is supporting or starring.
How much time should audio take in a typical project?
Budget 20 to 30 percent of total production time for audio. Teams that treat audio as a final five-minute step are the same teams that end up re-cutting the video to fix a narration that never fit.



