Most AI video projects do not fail because the generation model is weak. They fail because nobody translated the story into a set of shots before hitting generate. A writer has a scene in their head, types a paragraph into a text box, gets a beautiful but generic clip, and then spends the next two hours trying to bend it toward the story they actually meant to tell.
The fix is unglamorous: treat the screenplay as the source of truth, then convert it into a structured shot list before any pixels are rendered. That conversion step — script to beats, beats to shots, shots to prompts, prompts to reviewable clips — is where quality is won or lost. This guide walks through that pipeline in detail, with practical decision criteria, worked examples, and the mistakes that cost the most time.
Why the Script Still Wins in AI Video Production
Generative video has made individual shots cheap. It has not made coherent storytelling cheap. An audience will forgive a slightly soft frame or an odd hand; they will not forgive a scene that has no point of view, no escalation, and no visual logic from cut to cut.
A screenplay does three jobs that a prompt cannot do on its own. It establishes intent (what the scene must accomplish), it establishes sequence (what must be true before the next thing happens), and it establishes constraint (what must not change). When you skip straight to prompting, you lose all three. The model fills the vacuum with whatever is statistically plausible, which is why so much generated footage looks simultaneously impressive and interchangeable.
The practical consequence: the earlier you lock the script, the cheaper everything downstream becomes. Rewriting a line of action in a text document takes seconds. Re-generating twelve clips because the location drifted takes hours.
The Core Pipeline: From Screenplay to Shot List
Think of the workflow as four narrow gates. Each gate has a single job, and nothing passes through until that job is done.
Step 1 — Beat extraction and scene intent
Read the scene and write down its beats: the discrete turns in the action or emotion. A thirty-second scene usually has two to four beats. Label each one with a verb phrase — "she notices the door is open," "he decides to lie," "the lights cut out." If you cannot name the beat, it is probably not a beat; it is atmosphere, and atmosphere belongs in the shot description, not the beat list.
For each beat, record the intent in one line. Intent is what the audience should feel or understand, not what the camera sees. "Establish that the offer is a trap" is intent. "Close-up on a contract" is execution. Keeping these separate prevents you from locking in a visual solution before you have explored cheaper ones.
Step 2 — Shot list construction
Convert beats into shots. One beat rarely equals one shot, and one shot rarely equals one beat — a single beat may need a wide to orient the viewer plus a close-up to land the emotion, while a long take may cover three beats without a cut.
A workable shot list row contains: shot number, beat reference, shot size, camera movement, subject and action, location, time of day, lighting note, duration estimate, and continuity locks. That last column is the one people forget and later regret. Continuity locks are the fixed facts the shot must share with its neighbours: wardrobe item, prop in hand, hair state, weather, and the direction the light comes from.
Step 3 — Prompt translation and reference locking
Only now do you write generation prompts. Each prompt should contain five ingredients in a consistent order: subject, action, environment, camera, and style references. Consistency of ordering matters more than elegance of phrasing — models respond well to predictable structure, and you will catch omissions faster when the template never changes.
Where the tooling supports it, attach the same reference image for a character across every shot in a scene, and the same location plate for every shot in a location. Reference locking is the single highest-leverage habit in the entire pipeline.
Step 4 — Review against the beat sheet, not against taste
When reviewing output, ask first whether the shot serves its beat. A gorgeous clip that muddies the beat is a failure. Only after the beat check do you evaluate technical quality. This ordering stops the common spiral of polishing shots that should be cut entirely.
Writing Beats That Survive the Translation Into Prompts
Not all prose survives contact with a video model. Adjusting your writing style at the script stage saves enormous rework later.
Write visible action, not internal states
"She realizes she has been betrayed" cannot be rendered. "Her hand stops halfway to the glass, and she sets it down untouched" can. Every time you write an internal state, immediately write the external behaviour that proves it — and keep the behaviour in the shot list.
Keep location and light stable within a scene
Writers love to move characters around a building for realism. Video models treat every location change as a fresh visual problem. Consolidate locations where the story allows it, and when you must move, define the spatial relationship between areas explicitly so the audience never loses orientation.
Name the subtext in a note, not in the prompt
Subtext informs performance and framing choices. Putting it in the prompt invites the model to literalize it. Keep a separate director's note per scene — a sentence on tone, pace, and what the audience should feel — and translate that note into concrete camera and lighting decisions in the shot list.
Designing Shots: Framing, Movement, and Lens Language
Shot design is where AI video stops looking like a demo and starts looking like film. You do not need a cinematography degree, but you do need a small, consistent vocabulary.
Shot size vocabulary
Work with five sizes and resist adding more: wide (geography and isolation), medium-wide (body language and blocking), medium (conversation standard), close-up (emotion and detail), and extreme close-up (texture, tension, insert). Most scenes can be covered with three of these. When a scene feels flat, the problem is usually that every shot is the same size.
Camera movement that AI handles well
Slow push-ins, lateral tracking, gentle handheld drift, and static locked-off frames are reliable. Fast whips, complex arcs, and moves that require precise subject tracking under occlusion are still risky. Design your visual grammar around the reliable set rather than fighting the unreliable one, and save ambitious moves for the one or two hero shots where a few extra attempts are worth it.
A useful rule: movement should have a motivation. If you cannot say what the move reveals or denies, cut it. Gratuitous motion is the fastest way to make generated footage feel artificial.
Light and color as continuity anchors
Pick a key light direction, a colour temperature, and a contrast level per location, and repeat them in every shot there. Audiences read consistent lighting as spatial continuity even when the frame changes completely. Varying the light without story reason reads as a mistake.
Use colour to track emotional or narrative state deliberately. A scene that begins warm and ends cool tells the audience something without a line of dialogue. Just make the shift gradual and motivated, not a hard jump between adjacent shots.
Consistency Across Shots: Characters, Props, and Wardrobe
The most common complaint about AI-generated sequences is that the character changes between cuts. The problem is rarely model capability alone; it is the absence of a locking system.
Reference images and character sheets
Build a character sheet before you generate scene footage: one clean, evenly lit portrait plus three-quarter and profile views, plus a full-body shot in the primary costume. Generate this sheet once, review it carefully, and then use it as the visual anchor for every shot in which that character appears. Fix the sheet before you fix the shots — a weak anchor propagates everywhere.
Wardrobe and prop lock lists
Write down, per scene, exactly what each character wears and carries. Then check the shot list against that list. Small drifts — a jacket that changes shade, a bag that switches shoulders — read as continuity errors even to viewers who cannot articulate why something feels off.
When to break consistency on purpose
Deliberate change is a storytelling tool. A character who removes a coat after a confrontation, or whose hair comes loose over a long night, communicates time and pressure. Mark these changes as intentional in the shot list so you do not accidentally "fix" them during review, and make them happen at a cut rather than mid-shot.
Pacing, Rhythm, and the Assembly Problem
Editing is where most AI video projects quietly fall apart. Individually good clips assembled in the order they happened to be generated rarely produce a coherent rhythm.
Start by estimating durations at the shot-list stage. Short shots (one to three seconds) accelerate; long shots (five to eight seconds) decelerate. A scene with eight shots all cut to the same length will feel mechanical no matter how good the imagery is.
Then build an assembly pass before you polish anything. Lay the clips in order, watch once without stopping, and note only three things: where attention drops, where geography becomes confusing, and where the emotional turn fails to land. Fix those with reordering, trims, and the occasional inserted shot. Do not colour grade, do not add music refinement, and do not re-generate for aesthetic reasons during this pass.
Sound design deserves an early pass too. Room tone across a scene, a consistent ambient bed for a location, and deliberate use of silence will hide a surprising amount of visual inconsistency and will make a sequence feel intentional.
Matching Shot Types to Generation Approaches
Different shots reward different production strategies. A rough decision framework:
- Establishing and wide shots — generate several variations and choose on composition. These are forgiving of detail and easy to redo.
- Dialogue mediums — generate with a locked camera and a character reference. Prioritise facial stability over dramatic angle.
- Close-ups with emotional beats — spend the most attempts here. Small performance differences matter enormously, and a great close-up can carry a weak wide shot, never the reverse.
- Insert and texture shots — often the cheapest way to patch continuity gaps in the edit. Keep a small library of them per location.
- Complex action or effects shots — reduce ambition in the shot list, then compensate in the edit with cutaways. Two simpler shots almost always beat one over-specified complicated one.
A complementary option is to generate still keyframes first, approve the visual design across the whole scene, and animate from those keyframes. This front-loads the design decisions and dramatically reduces wasted motion generation.
A Practical End-to-End Workflow
Here is the sequence in the order it should be executed for a short narrative piece of roughly ninety seconds.
- Write the script — two to three pages, scenes kept short, locations consolidated, action written as visible behaviour.
- Extract beats — list every turn in each scene with a one-line intent.
- Build the shot list — sizes, movements, durations, continuity locks, and beat references.
- Create anchors — character sheets for each lead, location plates for each setting, a lighting note per location.
- Generate a rough pass — one attempt per shot, no polishing, accept imperfection to get a complete assembly.
- Assemble and review — fix structure first: order, duration, missing coverage.
- Re-generate selectively — only the shots whose beats fail, using the anchors and a tightened prompt template.
- Lock picture, then sound — add room tone, music, and any dialogue or narration.
- Final consistency sweep — scan for wardrobe, prop, and lighting drift that survives the edit.
The rough-pass discipline in step five is the hardest to accept and the most valuable. Projects that fully polish shot one before generating shot twenty routinely discover at assembly that shot one is unnecessary.
Common Mistakes and How to Avoid Them
Prompting the story instead of the shot. A paragraph of narrative in a prompt produces a montage of whatever the model finds plausible. Prompt one shot at a time.
Changing too many variables between attempts. If a shot fails, change one thing — camera, or wording, or reference — not all three. Otherwise you learn nothing about why the successful version worked.
No duration discipline. Generating long clips and trimming later feels flexible but wrecks rhythm. Decide the intended duration before generating.
Treating references as optional. Consistency problems are almost always anchor problems. Build the sheets and plates first.
Skipping the beat check. Reviewing on aesthetics rather than story function leads to beautiful footage that does not connect.
Over-specifying effects shots. The more a prompt demands simultaneous precision — specific subject, specific motion, specific environment — the more likely the result collapses. Split the demand across two shots.
Polishing before structure. Grading and refining a sequence that will be reordered is wasted effort.
Ignoring sound until the end. Ambient consistency is a continuity tool, not a finishing touch.
Frequently Asked Questions
Do I need to be a screenwriter to make good AI video?
You need to think in beats, not necessarily in screenplay format. A page of numbered beats with intents is enough structure to build a shot list from. Formatting conventions help collaboration but are not the source of quality — intent and sequence are.
How long should a shot list be for a one-minute video?
As a rough guide, one minute of finished narrative video typically uses twelve to twenty-five shots, depending on pace. Action and montage sequences skew higher; dialogue and mood pieces skew lower.
Why do my characters keep changing between shots?
Usually because each prompt was written independently. Fix it by creating an approved character sheet first, attaching the same reference wherever the tool allows, and keeping the wardrobe description identical, word for word, across every prompt in the scene.
Should I generate video directly or animate from keyframes?
If visual design is the priority and motion is simple, keyframes first gives you more control and fewer wasted attempts. If motion and timing are the point, direct generation is faster. Many projects use both — keyframes for hero shots, direct generation for connective shots.
How many attempts should I allow per shot?
Budget two to four for ordinary shots and six or more for emotional close-ups or complex movement. If a shot consistently fails after that, the problem is the shot design, not the model. Simplify it or cover the same beat with two easier shots.
Does shot order matter during generation?
Not technically, but generating in scene order keeps continuity fresh in your mind and lets you reuse environment descriptions and anchors with fewer transcription errors. It also makes assembly faster because the clips are already named and grouped sensibly.
What is the single biggest quality upgrade?
Adding a continuity-lock column to your shot list. It is a small documentation habit that prevents the majority of the drift, mismatch, and re-generation work that makes AI video projects feel slow.
The pattern across all of these answers is the same: the model handles rendering, and you handle decisions. Script, beats, shots, anchors, review. Do those in order and the generation step becomes the easy part.




