Sound Is Half the Story, and Most Editors Treat It Last
Ask any viewer what made a scene feel cinematic and they will usually describe what they saw. Ask them to watch the same scene with the music stripped out and they will describe it as flat, slow, or unfinished. That gap is where audio lives: it does not compete with the image, it tells the image what to feel.
This is why generative audio tools have become one of the most practical additions to a video workflow. Writing a custom score used to require either a composer, a licensed library, or hours of hunting through stock tracks that almost fit. Now you can generate a cue that matches a 42-second sequence, produce layered ambience for a location you have never recorded, and audition three emotional directions before lunch.
The catch is that AI audio does not remove the craft. It relocates it. Instead of playing an instrument, you are making decisions: where a cue enters, how loud a room tone sits under dialogue, when silence is more powerful than a build, and how a generated stem survives a final mix. This guide walks through a full workflow for music and sound design in an AI-assisted video pipeline, with the decision criteria, prompting technique, mixing targets, and common mistakes that separate audio that supports a story from audio that fights it.
What AI Audio Tools Actually Do Well
Before you build a workflow, be honest about the division of labor. Generative audio is excellent at some jobs and mediocre at others, and knowing the boundary saves hours of fighting the tool.
Strong fits for generation
- Mood-based background score. Ambient beds, tension pulses, lo-fi loops, corporate uplift, minimal piano — anything where the job is to establish tone rather than carry a melody the audience will hum.
- Sound effects and textures. Whooshes, risers, impacts, mechanical hums, crowd murmur, weather, sci-fi interface blips.
- Room tone and atmosphere. Adding a believable floor of sound to a space that was shot silently or recorded in a dead room.
- Alternate versions. Generating a 15-second cutdown, a 30-second variant, and a full-length cue from the same brief without re-licensing anything.
- Scratch tracks. Rough music to cut against so you can judge pacing before you commission or finalize anything.
Weak fits for generation
- Precise hit points. If a cue must land a downbeat exactly on a door slam at frame 1148, generate the cue without the hit and add the impact yourself as a separate element.
- Signature themes. A recurring three-note motif that identifies a character is possible to generate, but hard to reproduce consistently across many generations.
- Vocal-led hooks. Sung lyrics in a specific language, with a specific accent, and with semantically accurate phrasing remain unreliable.
- Dialogue replacement. Synthetic speech is strong for narration and scratch dialogue, but emotional performance, overlapping conversation, and mic-consistent continuity are still editor territory.
A useful rule: let AI produce the material, and let you produce the structure.
The Spotting Pass: Deciding Where Sound Belongs
Spotting is the traditional term for sitting with the picture and marking where music and effects should go. It is the highest-leverage 20 minutes in the entire audio process, and it is the step most people skip.
Mark the emotional beats first, not the timecodes
Watch your cut once with the sound off and no notes. Just notice where you feel something. Then watch again and write down the emotion and the approximate moment:
- Intro curiosity — establish tone, low commitment, sparse texture.
- Setup calm — background bed, low energy, does not distract from dialogue.
- Rising doubt — start introducing a pulse or a sustained note.
- Turn or reveal — the biggest change in the piece. Music either stops or peaks.
- Resolution — a chord or texture that resolves the tension you built.
- Outro — let the tail breathe, then exit cleanly.
Write these as plain-language briefs before you open any tool. "Sparse, curious, slightly cold, no drums, ends on an unresolved note." That single sentence is worth more than any preset list.
Convert beats into in and out points
Now map each emotional beat to actual timecodes. Decide three things per cue:
- In point: ideally on a cut, a movement, or a breath — not mid-sentence.
- Out point: usually a hard cut to silence or a fade under the last line of dialogue.
- Overlap: whether the cue should cross a scene transition or stop at the boundary. Overlapping cues create continuity; hard stops create contrast. Choose deliberately.
A practical constraint: for a two- to five-minute piece, three to five music cues is plenty. More than that and the audio starts feeling like a playlist rather than a score.
Prompting Music and Effects Like an Editor, Not a Poet
Generative audio responds best to structured descriptors. The most common failure is a prompt full of adjectives with no production language, which produces something generic that technically matches the mood and fits nothing.
The five-slot music prompt
Build every music prompt from five slots, in this order:
- Function — what the music must do. "Background underscore under narration, must not compete with speech."
- Genre or instrumentation — "solo felt piano, soft cello pad, brushed percussion."
- Texture and register — "low-mid heavy, no bright cymbals, narrow stereo image."
- Energy curve — "starts sparse, adds a pulse at the midpoint, pulls back to nothing at the end."
- Constraints — "no vocals, no lead melody, loops cleanly, 60 seconds."
Example: Background underscore under narration. Solo felt piano with a soft string pad, brushed percussion entering late. Warm low-mid texture, no bright high frequencies. Sparse at the start, builds gently, returns to near silence at the end. No vocals, no strong melodic hook, clean loop, about 60 seconds.
That prompt gives the model a shape, not just a vibe — and shape is what makes generated audio editable.
The three-part sound effect prompt
For effects, describe source, action, and space:
- Source: what is making the sound. "Heavy wooden door, metal latch."
- Action: how it moves. "Pushed open slowly, then a firm stop."
- Space: the acoustic environment. "Small stone room, close mic, some natural reverb tail."
Add duration and whether you need a clean tail or a hard cut. Requests like "one second, dry, no reverb" are far easier to place in a timeline than an atmospheric wash you then have to trim.
Generate in layers, not in one giant request
Do not ask for "cinematic trailer sound design for a product film." You will get a mush of everything. Generate separately:
- Music bed
- Low-frequency impact or sub swell
- Transition whoosh
- Ambience or room tone
- Accent effects on specific actions
Layered generation means you can mute what you hate without losing the rest, and it keeps your mix options open at the end of the project.
The Three-Layer Approach to Sound Design
Nearly every professional scene is built from three audio layers working together. Reproducing that structure with generated material is the fastest route to audio that sounds finished.
Layer 1: Ambience or room tone
This is the continuous floor of sound that tells the audience where they are. A cafe, a server room, a wind-swept field, a quiet apartment at night. Ambience is usually the quietest layer in the mix — often 15 to 25 dB below dialogue — but its absence is instantly noticeable. If a scene sounds "empty" or "cheap," missing ambience is almost always the reason.
Generate ambience in long, loopable chunks of 30 to 60 seconds and crossfade two copies for seamless coverage across a scene.
Layer 2: Hard effects and accents
These are the punctuating sounds tied to on-screen action: footsteps, a cup landing, a door, a keyboard, a car passing. They give the image physical weight. Generate these individually and place them frame-accurately. Slight timing offsets — a frame or two ahead of the visual — often read as more natural than a perfectly locked hit.
Layer 3: Music
Music carries emotion and pacing. It sits between the ambience and the accents in terms of prominence, and it should be the layer you are most willing to cut. If a scene works without music, music will only make it better. If a scene only works because of music, you have a structural problem worth fixing in the edit.
How the layers interact
| Layer | Typical role | Level relative to dialogue | Generated in |
|---|---|---|---|
| Ambience | Place and continuity | Very low | Long loops |
| Hard effects | Physical realism | Medium, brief | Individual hits |
| Music | Emotion and pacing | Low to medium | Full cues |
Once you think in layers, mix problems become diagnosable. Dialogue muddy? Check whether ambience is masking the 200–500 Hz range. Scene feels weightless? Add hard effects before adding music. Emotion lands wrong? The music cue is probably doing too much too early.
Dialogue, Narration, and Voice in an AI-Assisted Pipeline
Speech is the most demanding element of any mix because the audience is processing meaning, not just sound. Treat voice as the anchor and everything else as support.
When to generate voice
Synthetic narration works well for explainers, internal training, documentary scratch tracks, and multilingual versions of a script you already approved. It works less well when the performance itself is the content — comedy, personal storytelling, character acting.
Practical rules for generated voice
- Write for the ear, not the page. Short sentences, natural contractions, one idea per line.
- Generate paragraph by paragraph. Long generations drift in energy and pacing. Short chunks give you consistent tone and easy retakes.
- Keep a pronunciation list. Product names, acronyms, and place names should be written phonetically in your script once and reused everywhere.
- Leave headroom between lines. Natural pauses are easier to shorten than to invent.
- Match room tone. If narration must sit inside a scene, add a touch of the same ambience under it so it does not sound pasted on.
Balancing voice against generated music
The single most effective technique is a duck: lower the music by 4 to 8 dB whenever dialogue is present, with a smooth attack of 100–300 ms and a release of 300–600 ms. Most editors can do this with a sidechain compressor or keyframed volume automation. If you are mixing on a phone editor, simply cut the music level manually at each line — the result is the same and the intent is clearer.
Mixing and Loudness: Making AI Audio Delivery-Ready
Generated audio arrives at wildly inconsistent levels. Normalization is not optional; it is the step that makes your project sound intentional.
A workable starting balance
- Dialogue: the loudest sustained element, peaking around -6 dBFS.
- Music: 12 to 18 dB below dialogue under speech, rising to 6 to 10 dB below in music-only passages.
- Ambience: 20 to 30 dB below dialogue.
- Hard effects: brief peaks near dialogue level, never sustained.
Loudness targets by destination
| Destination | Integrated loudness | True peak ceiling |
|---|---|---|
| Web video and streaming | about -14 LUFS | -1 dBTP |
| Social short-form | -14 to -12 LUFS | -1 dBTP |
| Podcast and spoken audio | -16 LUFS | -1 dBTP |
| Broadcast delivery | per the broadcaster's spec | -2 dBTP |
Measure with a loudness meter, not by ear alone. Ears adapt within minutes; meters do not.
Clean the artifacts before you mix
Generated audio often carries small defects: a hum at 50 or 60 Hz, a hiss above 12 kHz, an abrupt loop point, or a clipped tail. Run a high-pass filter at 30–40 Hz on music and effects, notch any obvious hum, and always add short fades of 5–15 ms at the start and end of every clip. These three habits eliminate most "why does this sound amateur?" problems.
Check the mono fold-down
A large share of viewers watch on a phone speaker. Sum your mix to mono and listen. If music or ambience swallows the dialogue in mono, the wide stereo version is hiding a problem you should fix in stereo.
A Repeatable End-to-End Workflow
Here is the sequence that keeps projects predictable, whether you are cutting a 30-second ad or a 12-minute documentary.
- Lock picture. Do not score a moving target. Minor timing changes after scoring mean re-doing everything downstream.
- Spot the timeline. Write emotional beats in plain language and map them to in and out points.
- Lay ambience first. Cover every scene with room tone before adding anything else.
- Add hard effects. Place accents on action, letting slight offsets feel natural.
- Generate music cues. One cue per emotional beat, prompted with the five-slot structure.
- Scratch-mix quickly. Get rough levels in 15 minutes and watch the whole piece once without stopping. Judge the story, not the details.
- Refine and duck. Automate music under dialogue, fix masking, trim unnecessary cues.
- Clean and normalize. Filters, fades, loudness measurement, true peak ceiling.
- Check on real devices. Phone speaker, laptop, headphones. Three checks catch nearly everything.
- Export stems and archive. Keep music, ambience, effects, and dialogue as separate files so future edits do not require regeneration.
Steps 1, 2, and 6 are the ones people skip, and they are the ones that determine whether the finished piece sounds deliberate.
Common Mistakes and How to Fix Them
Generating music before spotting. You end up shaping the story around the track instead of the track around the story. Fix: reverse the order, even if it feels slower at first.
One giant prompt for everything. Results are muddy and impossible to adjust. Fix: layer music, ambience, and effects as separate generations.
Music that never stops. Constant score flattens the emotional range of an entire piece. Fix: cut the music out for 5 to 10 seconds before your biggest moment. Silence is the cheapest and most powerful transition available.
Ignoring ambience. The classic symptom of AI-heavy edits: scenes feel like they are happening in a vacuum. Fix: bed every scene with room tone, even a very quiet one.
Over-loud effects. A single whoosh at dialogue level will make a viewer flinch. Fix: keep accents brief, and check the mix on a phone speaker.
Stylistic drift between cues. Three cues generated separately can sound like three different films. Fix: reuse the same descriptive vocabulary and instrumentation list across every prompt in a project, and pin your key elements (tempo range, instrumentation, register) into every request.
No loop discipline. Ambient beds with clicks or gaps at the seams break immersion. Fix: ask for clean loops, then crossfade two copies over 1–2 seconds when placing them.
Skipping the archive. If you overwrite or discard the individual stems, your next revision is a full regeneration. Fix: export and name stems by scene and layer.
Frequently Asked Questions
Do I still need a composer or sound designer?
For high-stakes brand films, narrative shorts, and anything where audio is a headline feature, yes — a human will get you to a result faster and with more control. For the large middle ground of explainers, social content, internal video, and documentation, a structured AI workflow covers the job well.
How long should a music cue be?
As long as the emotional beat it supports, plus two seconds of tail. A 20-second beat does not need a 90-second track; trimming a long cue usually leaves an awkward ending.
Can I mix generated audio with licensed library tracks?
Yes, and it is often the smartest approach. Use library music where you need a proven hook and generated audio for transitions, textures, and scene-specific beds that no library track fits.
How do I keep a series sounding consistent?
Treat your audio like a brand asset. Save one document with your instrumentation list, tempo range, ambience palette, and loudness target, and reuse it for every episode. Consistency in audio reads as professionalism even when viewers cannot name what they are hearing.
What is the fastest way to improve a weak mix?
Three moves: add ambience under every scene, cut music out before your biggest moment, and lower music by 6 dB under all dialogue. Most mixes improve dramatically with only those changes.
Should I generate one long track or several cues?
Several cues, always. Long tracks lock you into a single emotional register, and they are far harder to re-time when the edit changes.
Start With Structure, Finish With Restraint
AI music and sound design tools are genuinely transformative, but not because they replace taste. They compress the distance between an idea and a usable piece of audio, which means the bottleneck moves to judgment: knowing where sound belongs, how loud it should be, and when it should disappear.
If you take one habit from this guide, make it the spotting pass. Twenty minutes spent marking emotional beats turns a random collection of generated clips into a score. Everything after that — prompting in layers, ducking under dialogue, cleaning loop points, hitting a loudness target — is mechanical work that gets faster with repetition.
Build the workflow once, document your instrumentation and ambience palette, and archive your stems. Then the next project starts with a system instead of a blank timeline, and your audio stops being the last thing you fix and becomes the reason the piece lands.




