The real bottleneck is intention, not generation
Anyone with a browser can now turn a sentence into a moving image. What almost nobody can do reliably is produce the right moving image — the one that carries a story, matches the shot before it, and lands the emotional beat you planned three scenes ago. That gap between "a video came out" and "the video I needed came out" is where narrative prompt design lives.
Treat prompts as a craft layer sitting between your script and your timeline. The creators who ship coherent short films, explainers, and episodic series with generative video are not using secret models. They write clearer instructions in a consistent format, and they do the boring continuity work most people skip.
This guide covers the anatomy of a narrative prompt, how to keep characters and props stable across dozens of shots, how to pace emotion with shot length, how to combine stills and audio with text, and a repeatable workflow you can run at any project length.
Anatomy of a narrative prompt that holds up
A descriptive prompt says what is in frame. A narrative prompt says what is happening, to whom, why it matters right now, and how the camera should feel about it. In practice, the strongest prompts are assembled from five blocks, always in the same order, so you can debug one block at a time instead of rewriting everything.
Block one: the character
Name the subject, give one or two physical anchors, state the emotional state, and specify wardrobe. "Maya, mid-thirties, close-cropped black hair, charcoal field jacket, jaw tight with restrained anger" outperforms "a woman who is upset" every time. Emotion adjectives alone are noise to a diffusion model; physical cues that read as emotion are signal.
Block two: the world
Describe location, time of day, weather, and dominant palette. Two or three concrete details beat ten vague ones. If the scene is a rain-slick alley at dusk, name the light sources — a sodium streetlamp, a flickering neon sign, a doorway spill — because light sources determine how the subject is lit, and lighting is the fastest route to visual continuity between shots.
Block three: the beat
This is the most neglected block. A beat is one action with a beginning and an end: she opens the envelope, reads it, looks up. Not "she discovers the truth." Prompt one beat per generated shot. If you cram three beats into one clip, the model rushes all three and finishes none, and you cannot cut around the result.
Block four: camera
Specify shot size, angle, lens feel, and movement. "Medium close-up, eye level, 50mm feel, slow push in" gives you a shot you can match to its neighbor. "Cinematic" gives you nothing. Consistent camera language is what makes a sequence of clips feel shot by one crew rather than assembled from five different projects.
Block five: sound and speech
If the tool supports audio, describe ambience, music direction, and any spoken line with its tone of delivery. Even for silent-only tools, writing sound design into the prompt keeps your shot lengths honest — a line of dialogue has a natural duration, and that duration should determine how long the clip needs to be.
Build the story spine before you generate anything
The most common failure mode in AI video work is generating before thinking. Ten beautiful clips that do not cohere are worth less than four average clips that do. Before you open a generator, produce four small artifacts.
The logline. One sentence: who wants what, what blocks them, what is at stake. If you cannot write it, the piece is not ready.
The beat sheet. Eight to twelve beats for a short piece. Each beat gets a one-line description and an emotional value — tension, relief, curiosity, dread.
The shot list. Each beat becomes one to three shots. For each shot, note size, subject, action, and target duration.
The continuity bible. A plain text document listing every recurring character, prop, and location with fixed descriptors. This is what you copy from every time you write a prompt, and it is why your fifth shot looks like your first.
That is twenty to forty minutes of prep for a ninety-second film. It saves hours of regeneration.
Keeping characters and objects consistent
Consistency is the hardest technical problem in AI video, and no tool solves it for you. You solve it with references, locked descriptors, and discipline.
Reference images and first-frame locking
Generate a clean, neutral reference of each character — front-facing, even lighting, mid-shot, no dramatic expression. Use it wherever the tool accepts image conditioning. Where the tool supports a first-frame or keyframe input, feed the final frame of the previous shot in as the starting frame of the next. This turns continuity from a hope into a constraint.
Wardrobe, props, and fixed strings
Write descriptors once, exactly, and paste them verbatim. The moment you paraphrase — "grey coat" in one prompt, "gray overcoat" in the next — you introduce variation. Recurring props deserve the same treatment: the brass key, the red notebook, the cracked phone screen. Give each a short fixed string such as brass key, worn edges, warm highlight and never improvise it. Alternating between "the woman," "she," and "Maya" across shots is another quiet source of drift; pick one label and keep it.
Accept controlled drift
Perfect identity lock across a long sequence is still rare. Design around it: use more over-the-shoulder shots, silhouettes, and hands-in-frame inserts, and save clean frontal close-ups for moments where the audience needs the face. A story that limits exposure to the hardest shots looks intentional rather than glitchy.
Pacing: mapping emotion to shot length
Shot duration is a narrative instrument. A three-second clip that cuts early creates momentum; a six-second clip that holds creates unease or intimacy. If you generate everything at the default duration, you forfeit that instrument.
A practical default rhythm for narrative work:
- Establishing shots: four to six seconds, slow movement, no internal cuts.
- Dialogue or reaction shots: two to four seconds, matched to the spoken line.
- Action beats: one to two seconds, faster camera, cut on the movement.
- Emotional holds: five to eight seconds, minimal camera motion, let the face carry it.
- Transitions: half a second to a second of connective material — an insert, a hand, a door.
Match the clip length you request to the edit length you need, plus a beat of handle on each end. Generating a ten-second clip to use three seconds of it wastes time and tempts you into using footage that is too slow for the scene. The most reliable rule in the entire workflow: one idea per shot. If a clip contains two ideas, you cannot cut it without losing one.
Combining images, audio, and script
Text prompts on their own drift. Adding a second modality — a still image, a voice track, a reference clip — pins the output to something specific.
Stills as visual anchors
A still generated in an image tool and then animated gives you far more control than pure text-to-video. You choose framing, lighting, and wardrobe in a cheap, fast medium, then ask the video model to animate within those constraints. Build a small library of stills per scene and animate them in order.
Voice, ambience, and music
Record or synthesize dialogue first, then cut visuals to the audio rather than the reverse. This is how real production works, and it eliminates most lip-sync problems because you time the shot to the line instead of hoping the line fits the shot.
For ambience, layer three things: a bed (room tone, rain, traffic), a mid layer (footsteps, cloth, keyboard), and an accent (a single door slam, a phone buzz). Two of the three can be library audio; the accent sells the shot.
Practical lip-sync expectations
Frontal, well-lit, moderately paced dialogue is where sync tools perform best. Profile shots, extreme close-ups of the mouth, heavy accents, and rapid speech are where they fail. Write scenes where characters talk while doing something — walking, cooking, driving — so the face is not the only thing competing for attention.
A repeatable end-to-end workflow
One: write the spine. Logline, beat sheet, shot list, continuity bible.
Two: generate key art. One still per scene, plus neutral character references. Iterate here freely, because stills are cheap.
Three: lock the look. Pick palette, grain, and lighting treatment across the stills until they feel like one film.
Four: record audio. Dialogue, scratch voiceover, temp music. Now you know exact durations.
Five: generate shots in continuity order. Shot one, then two, feeding the previous final frame forward where possible. Do not jump around.
Six: assemble a rough cut with placeholder text. Get the rhythm right before the visuals are finished.
Seven: replace placeholders as shots arrive. Regenerate the worst offenders first.
Eight: sound design and grade. One grade, one ambience bed, one music cue across the whole piece.
Nine: publish and log what worked. Keep the prompts that produced usable shots so your next project starts from a library instead of a blank page.
Choosing a model without chasing hype
Ignore leaderboard rankings and test against your actual requirements. Five criteria matter more than raw fidelity: subject motion quality (does a walking person look like a walking person?), camera control (can you ask for a specific move and get it repeatedly?), usable clip duration before artifacts appear, whether the tool accepts a starting frame or character reference, and native audio support versus clean separation for post.
Test each candidate with the same three-shot sequence: a character walking into a room, a dialogue close-up, and a fast action beat. The tool that handles all three acceptably is the one you build a pipeline on, even if another tool wins on a single demo clip. Then consider a two-model pipeline — one for hero shots, one cheaper option for inserts and coverage. Most projects do not need every shot to be a showcase.
Mistakes that quietly ruin narrative coherence
Prompt creep. Changing wording between shots, then wondering why the face changed.
Multi-beat prompts. Asking for "she enters, sits, and cries" in one clip and getting three rushed half-actions.
Default durations. Accepting whatever length the tool returns instead of the length the edit needs.
Generating out of order. Starting with the fun shot instead of shot one, which breaks frame-to-frame continuity.
Describing mood instead of light. "Sad lighting" means nothing; "single window, overcast, cool grey, subject lit from camera left" means everything.
Visual-first editing. Building the cut before the audio forces awkward trims later.
No continuity bible. Relying on memory across a forty-shot sequence.
Over-polishing before assembly. Spending hours on one shot before discovering the sequence does not cut together.
Pre-publish quality checklist
- Every recurring character uses the identical descriptor string.
- Every shot contains exactly one narrative beat.
- Shot lengths vary with emotional intensity, not by accident.
- Camera language is consistent within each scene.
- Audio leads the edit; no clip is longer than the line it carries.
- The first three seconds establish who and where.
- The final shot resolves the beat established at the start.
- One grade and one ambience bed across the entire piece.
- Titles are legible on a phone before you export.
FAQ
How long should a prompt be? Long enough to hold all five blocks, short enough to read aloud in under thirty seconds. Eighty to a hundred and fifty words is a practical range; beyond that, models start dropping details.
Do I need a different format for every tool? No. Keep one master format and adapt the syntax. The blocks stay the same; only parameters change.
How do I fix a character whose face shifts between shots? Reintroduce the reference image, paste the identical descriptor string, and check that prompt length and phrasing match earlier shots. Then shoot around it with inserts and over-the-shoulder framing.
One long clip or many short ones? Many short ones. Long clips give you less control, drift more, and resist cutting. Think in shots, not scenes.
How do I handle dialogue-heavy scenes? Record audio first, animate to it, keep faces moderately lit and relatively frontal, and give characters physical business so the mouth is not the only thing competing for the viewer's attention.
What is the fastest way to improve results? Write the shot list before generating anything. Most quality problems are pre-production problems wearing a costume.
Can this scale to a series? Yes, and it accelerates. The continuity bible, still library, and prompt templates compound across episodes, so episode four costs a fraction of the effort of episode one.




