Why a Bespoke Score Outperforms a Stock Track
Library music is engineered to be inoffensive. The same track has to sit under a car commercial, a cooking montage, and a corporate explainer, so it avoids strong personality. That compromise is audible. A bespoke cue knows where the reveal lands, how long the closing card holds, and which line of dialogue it must never mask.
Commissioning that cue used to mean briefing a composer, waiting weeks, and negotiating usage terms. The threshold has dropped sharply. The distance between needing a very specific feeling under a shot and having six versions of it is now about ninety seconds of prompting, which changes the economics of scoring far more than it changes the craft.
Three practical consequences matter for anyone cutting video today:
- Iteration replaces commitment. Audition a warm analog bed, discover it fights the narration, and replace it with a sparse glassy pad in the same session.
- Motifs become reusable assets. Extract a short figure and re-orchestrate it per episode to build recognition across a series.
- Silence becomes a decision. Beginners score wall to wall; experienced editors leave air around the lines that matter.
None of this removes skill from the process. It relocates skill toward judgment: structuring emotion, describing sound precisely, choosing a take against picture, and mixing so the words always win.
How Text-to-Music Tools Actually Think
Most text-to-music systems are trained on large audio collections and learn statistical relationships between language and sound. Some generate spectrograms or waveforms through diffusion; others predict discrete audio tokens the way a language model predicts words. The machinery matters less than the consequence: the model renders acoustic description, not narrative intent.
Write sad scene and you receive a generic minor-key wash. Write solo cello, slow bowing, close microphone, reverb tail under two seconds, no percussion, sustained low drone entering halfway through and you get something you can place under a shot.
Behind almost any interface sits a similar pipeline: a text encoder parses the prompt, a generator produces an initial segment, a continuation mechanism extends it, and post-processing normalises loudness. Many tools then offer stem separation, which is what makes generated music genuinely editable — it turns a finished stereo file into rhythmic, bass, harmonic, and textural layers.
The landscape divides into four families:
- Song-oriented generators produce full arrangements with verses and choruses. Ideal for title sequences, montages, and anything needing a hook.
- Instrumental-first generators hold together over long durations and respond precisely to instrumentation prompts. Usually the right choice for underscore.
- Texture and sound-design tools create ambience, risers, impacts, room tone, and loops — the connective tissue between score and effects.
- Editor-side assistants cover stem splitting, noise reduction, auto-levelling, and mastering. Unglamorous, but they improve perceived quality more than any prompt.
A workable stack is one tool from each of the first three families, plus whatever helpers your editing software already includes.
Start with a Cue Sheet, Not a Prompt
Generating first and planning later is the fastest way to waste an afternoon. Walk the rough cut and mark not every cut, but every emotional shift. Then build a table:
| Timecode | Story beat | Desired feeling | Energy (1–10) |
|---|---|---|---|
| 00:00–00:12 | Cold open, drone over skyline | Unease, scale | 3 |
| 00:12–00:38 | Founder interview begins | Trust, warmth | 2 |
| 00:38–01:05 | Product montage | Momentum, optimism | 7 |
| 01:05–01:30 | Results and closing card | Resolve, lift | 4 |
That table is the score; everything after it is execution. It also tells you where you need silence, and where dense dialogue means the music should stay nearly invisible.
Translate each row into a paragraph a model can read, containing six ingredients: genre, instrumentation, tempo and key, energy curve, production texture, and how the piece should end. For example:
Sparse cinematic underscore, felted upright piano and low strings, 72 BPM, D minor, energy rising from 2 to 4 across the piece, close-miked with a soft small room, no drums, no vocals, no bright cymbals, no melodic resolution at the end.
Two habits improve results immediately. Describe references by attributes rather than names — reverb-drenched, wide stereo, breathy, tape-saturated — because attribute language transfers reliably between models. And always specify the ending, since generators default to a fade or an abrupt stop when you need a held note that hands off to the next scene.
Plan duration deliberately. Most systems produce thirty seconds to a few minutes per pass, so anything longer means extending a track or cutting between cues. Cutting between cues usually wins, because it gives you editorial control of the transition instead of hoping the model guessed.
Writing Prompts That Survive Contact with Picture
Prompt quality is the craft, and specificity is the only reliable lever. A strong prompt answers six questions at once: what tradition, what instruments, what register, what tempo, what texture, what shape.
Keep useful vocabulary nearby: pad, drone, ostinato, stinger, riser, sub drop, downbeat, half-time, tremolo, pizzicato, detuned, sidechained, lo-fi, tape-saturated, close-miked, plate reverb, long decay.
Build variation methodically. Generate six to twelve candidates per cue and change one variable at a time — hold the instrumentation and shift tempo, then hold tempo and swap the lead instrument. Name files so variables are recoverable, such as cue03_piano_72_v4, and note which variable each batch explored. Random regenerating destroys your ability to learn what worked.
Useful negatives: no vocals, no cymbals, no snare, no melodic resolution, no sudden stops, no busy arpeggios under dialogue. Avoid prompting for a living artist or a recognisable film score; beyond the rights problem, those prompts rarely yield publishable material. Describe the qualities you want instead.
The Scoring Workflow, Step by Step
1. Lock the picture. Never score a cut that still moves. Every tempo relationship breaks when a shot shortens by six frames.
2. Sketch a tempo map. Assign rough tempos per section and check they divide cleanly into scene durations.
3. Generate in batches. Six to twelve options per cue, one variable at a time, with disciplined naming.
4. Audit against picture. Audition with dialogue in place, never soloed. Music that sounds thin alone often sounds perfect under a voice, while lush tracks frequently bury words.
5. Split into stems. Separate the chosen track into rhythmic, bass, harmonic, and textural elements; you will almost always want one gone.
6. Edit to hit points. Mark moments the music must acknowledge — a logo landing, a punchline — and nudge a chord change onto the mark. Three or four deliberate hits per minute is plenty.
7. Mix under dialogue. Treat the score as support with a defined frequency pocket.
8. Master and archive. Deliver a full mix plus a music-and-effects stem, then store prompt, model name, date, and stems together. When the project returns in six months, that archive matters more than the audio itself.
Looping deserves a separate note. Generated loops rarely join cleanly. Find a point where the waveform crosses zero, cut there, and inspect the seam at low volume or in a spectrogram. If a reverb tail rings across the join, crossfade ten to forty milliseconds, or rebuild the loop from the rhythmic stem alone.
Tempo Math and Hit Points
The relationship between tempo and edit rhythm is arithmetic, and knowing it lets you plan instead of guess. At 120 BPM, one beat is half a second and a four-beat bar is two seconds. Halve the tempo and those durations double.
| Tempo | Beat length | 4/4 bar | Typical use |
|---|---|---|---|
| 60 BPM | 1.00 s | 4.0 s | Slow documentary underscore |
| 72 BPM | 0.83 s | 3.3 s | Reflective brand film |
| 90 BPM | 0.67 s | 2.7 s | Warm montage |
| 100 BPM | 0.60 s | 2.4 s | Product walkthrough |
| 120 BPM | 0.50 s | 2.0 s | Energetic launch |
| 140 BPM | 0.43 s | 1.7 s | Fast social edit |
Choose a tempo whose bar length divides neatly into your scene durations. A thirty-second montage at 120 BPM holds fifteen bars, giving clean two-bar and four-bar phrase endings. A twenty-eight-second cut will fight the grid the whole way. Hit points are deliberate exceptions: pick your three most important moments and align them precisely, then let everything else breathe. Audiences notice precision at emotional peaks and forgive looseness elsewhere.
Mixing the Score Beneath Dialogue
The score is one voice in a conversation, and it loses every fight it picks with dialogue, so it should not pick any.
Carve a pocket. Speech intelligibility lives roughly between 1 kHz and 4 kHz. Use gentle dynamic equalisation to pull two to four decibels from the music whenever a voice is present, and release it in the gaps.
Duck, do not crush. Broadband ducking of three to six decibels with slow attack and release sounds natural. Heavy sidechaining makes music pump audibly and falls apart on headphones.
Level with intention. Dialogue should sit clearly above the bed at all times. If you keep pushing music up to feel impactful, the real problem is arrangement density, not volume.
Let effects carry impact. Whooshes, risers, clicks, and sub drops hit harder than a loud musical stab and are trivial to generate with a texture tool. Reserve the score for continuity and feeling.
Check tiny speakers. Phone speakers are the real distribution format for short-form video. If the score nearly disappears there while dialogue still reads, you have balanced it correctly.
Matching Musical Texture to Visual Style
Visual grammar should drive musical texture. Mismatches read as amateur even when both halves are well made.
| Visual style | Musical treatment | Avoid |
|---|---|---|
| Slow aerials and drone footage | Sustained pads, low strings, minimal percussion | Busy rhythms, sharp transients |
| Handheld documentary | Acoustic instruments, room noise, imperfect timing | Over-quantised synths |
| High-contrast product macro | Tight percussion, sub bass, short reverb | Long washes that blur cuts |
| Archival or grainy footage | Tape saturation, vinyl noise, mono feel | Pristine modern mixes |
| Stylised animation | Playful orchestration, motifs per character | Wall-to-wall scoring with no air |
| Talking-head interviews | Nearly invisible pads, or nothing | Anything with a memorable melody |
One rule covers most cases: the more information the image carries, the less the music should say. A dense, fast-cut montage needs a simple, steady foundation; a static interview shot gives you room for harmonic movement. If your footage itself comes from generative video tools, keep a single aesthetic throughline — soft, filmic visuals paired with a harsh digital score feels assembled rather than designed.
Choosing a Tool and Vetting Its Output
Judge tools on concrete criteria rather than demo reels.
| Criterion | Why it matters | What good looks like |
|---|---|---|
| Control granularity | Determines whether problems are fixable by prompting | Section-level structure, tempo and key fields |
| Stem export | Enables editing without destructive regeneration | Clean four to six stem separation |
| Duration handling | Affects how much stitching you must do | Extend, continue, and out-painting support |
| Usage terms | Governs where you can publish | Plain-language, clearly stated rights |
| Batch generation | Speeds up exploration | Multiple variations per request |
| Offline options | Matters for confidential material | Local or self-hosted models |
For video work, prioritise stem export and control granularity over raw fidelity. Fidelity is improving everywhere; editability is what makes a tool usable on a deadline.
Before a score leaves your session, run this check: listen once at low volume so nothing pokes out; listen on phone speakers for dialogue legibility; check the final ten seconds, where generative tracks reveal themselves; verify the loop seam if the track repeats; confirm no clipping in the loudest passage; and confirm the mix survives mono playback and platform compression.
Treat usage terms as a production step, not an afterthought. Read the terms for your specific tool, keep a record of which model produced which cue, and export dated project files. Avoid generating anything that imitates a living artist or a signature sound without permission. Some platforms require disclosure when audio is synthetic, so check the rules of the channels you publish to and, when in doubt, state the process in your description.
The most common mistakes repeat across projects: scoring every second instead of letting silence work; generating before mapping the emotional arc; choosing a track in solo rather than under dialogue; ignoring tempo arithmetic and then fighting the grid; and never splitting into stems, which forces an all-or-nothing choice between a good track and a good mix. Each has the same fix — decide the structure on paper first, then use generation to execute it.
Frequently Asked Questions
Do I need musical training to do this well?
No, but you need vocabulary. Learn what a pad, stinger, riser, and downbeat are, then practise describing tempo and mood numerically. That is a weekend of focused listening.
Can one generated track carry an entire series?
Yes, and consistency builds recognition. Extract a short motif from your main cue and re-orchestrate it for different episodes or chapters.
How long should a single cue be?
Only as long as the emotional beat it supports. Most cues run twenty to ninety seconds, and longer is rarely better.
Why does my score sound generic?
Usually because the prompt is generic. Add instrumentation, register, articulation, microphone distance, reverb size, and explicit negatives.
Should I generate music or license a library track?
Generate when the music must fit a specific arc, tempo, or duration. License when you need a recognisable genre sound quickly and the picture can flex around it.
What about vocals?
Treat them as a separate element. Wordless or non-English textures are safer beneath dialogue; recognisable lyrics compete directly with your speaker.
The Bottom Line
Generative audio did not remove the skill from scoring — it moved the skill. The work now lives in structuring emotion, writing precise descriptions, choosing the right take against picture, editing to hit points, and mixing so the words always win. Generate generously, decide ruthlessly, and keep your project files. Months later, the prompt will be worth more than the audio it produced.



