Most disappointing AI video does not fail because the prompt was too short. It fails because nobody decided what the scene was about before the first frame was generated. Story structure and shot design are the two disciplines that turn a folder of clips into a film, and both are learnable through a repeatable process. This guide walks through that process end to end: translate a script into beats, convert beats into a shot list, and only then select generative models and write prompts. The goal is directorial control — the ability to predict how a shot will feel before you spend time rendering it.
Why Story Structure Must Come Before the Prompt
A prompt is an instruction. A story is an intention. When you write prompts first, you are asking a model to solve an artistic problem you have not fully defined yet. The result is usually technically impressive and emotionally flat: beautiful lighting with no dramatic reason to exist, camera moves that show off rather than reveal.
Begin with the dramatic question. Every scene should be able to answer three things in one sentence each: Who wants what, what is blocking them, and what changes by the end of the scene? If you cannot answer those, no model choice will rescue the footage.
Structure also determines what you are allowed to cut. Sequences that carry information need coverage — a wide to establish, a medium to humanize, a close-up to land the emotion. Sequences that carry feeling need fewer, longer shots. Knowing which is which before generation prevents the classic trap of generating fifteen clips and discovering that four of them say the same thing.
Finally, structure protects you from the seduction of novelty. A generative tool will happily produce an elaborate transformation shot that has nothing to do with your story. A beat sheet tells you whether you actually need it or whether you are simply excited by what the model can do.
The three documents worth writing before anything else
A one-page logline and synopsis, a beat sheet of twelve to twenty beats, and a scene intent note of two or three lines per scene. None of these need to be polished. They need to exist so that every later decision has a reference point. When a shot feels wrong in the edit, you diagnose it against these documents rather than against taste alone.
The Story-to-Shot Translation Layer
This is where most creators skip a step. The jump from "Maya confronts her brother in the rain" to a prompt is too large. Insert a translation layer: the shot table.
A shot table is a simple spreadsheet with one row per shot and columns for scene, beat, dramatic function, subject, framing, camera movement, duration, lighting mood, continuity notes, and generation notes. It is unglamorous and it is the single highest-leverage artifact in an AI video pipeline.
Breaking a scene into dramatic beats
Take a scene and mark where the emotional temperature changes. A conversation might have four beats: greeting, deflection, accusation, admission. Each beat gets at least one shot. Beats that carry the emotional turn deserve the closest coverage; connective beats can be covered with a single moving shot.
Assigning dramatic function to each shot
Every shot should have a job description. Useful labels include establish, orient, reveal, react, escalate, withhold, resolve, and transition. If a shot cannot be labeled, cut it from the plan. This single rule eliminates most filler.
Deciding coverage ratios early
Coverage ratio is the number of generated shots you plan against the number of shots you will use. For a dialogue scene, plan three to one. For a montage, plan two to one. For a hero visual effect shot, plan eight to one. Budgeting coverage in advance turns generation from a gamble into a schedule.
Avoiding the "cool shot" trap
Keep a separate parking lot for shots you love but cannot justify. Review the parking lot after the first assembly. Sometimes a parking-lot shot reveals a better structure; more often it saves you from spending half your render time on a moment that the story never needed.
Shot Design Fundamentals That Survive AI Generation
Shot design is the visual grammar that makes a scene readable. Generative models are now good enough that these fundamentals matter more than raw image quality, because audiences forgive soft detail and never forgive confusion.
Framing and lens language
Think in terms of wide, medium, close, and insert, and think about what each does to the audience's relationship with the character. Wide shots create context and isolation. Medium shots create conversation and neutrality. Close-ups create access to interiority. Inserts create specificity — the trembling hand, the unread message, the cracked phone screen.
Lens language is equally usable. Wide-angle framing exaggerates space and movement; longer-lens framing compresses depth and flatters faces. When writing prompts, describe the effect rather than the millimeter: "compressed background, shallow depth of field, subject isolated against a soft city bokeh" communicates more reliably than a technical focal length.
Camera movement vocabulary
Limit yourself to movements you can justify: slow push in for growing tension, pull out for revelation or isolation, lateral tracking for parallelism, handheld drift for unease, crane or rise for scale, static for observation. Generative models respond well to a single dominant movement per shot. Stacking a push, a pan, and a roll usually produces mush.
Blocking, eyelines, and screen direction
AI generation makes it easy to break the 180-degree rule without noticing, because each shot is produced independently. Track screen direction in your shot table: if a character moves left to right toward a door in the wide, they should continue left to right in the medium, or you must show the reversal on screen. Track eyelines too — who looks frame-left and who looks frame-right determines whether two people appear to be in the same room.
Light, color, and time of day as continuity anchors
Decide the lighting mood and color temperature for each scene once, and repeat that language in every prompt for that scene. "Overcast daylight, cool desaturated palette, wet asphalt reflections" used consistently will hold a sequence together even when faces drift slightly between shots.
Matching Model Choice to Shot Type and Intent
Different shot types stress generative systems in different ways. Rather than treating any single model as universal, build a small internal routing table.
Where each model family tends to excel
Text-to-video models with strong temporal coherence are best for continuous movement, environmental shots, and establishing wides. Image-to-video models are best when you already control composition and want motion added to a locked frame — ideal for close-ups and inserts. Motion-transfer and performance-driven tools are best for dialogue coverage, where facial nuance matters more than camera ambition. Upscaling and frame-interpolation utilities are not generation models at all but are essential for matching shots that came from different sources.
Decision criteria for routing a shot
Ask four questions: Does this shot need precise composition? Does it need complex motion? Does it need a recognizable face or product? Does it need to match a previously generated shot exactly? Precise composition and matching favor image-to-video. Complex motion favors text-to-video. Recognizable faces favor performance-driven approaches plus consistent reference images. If two answers conflict, split the shot into two: one that establishes, one that moves.
Speed versus fidelity budgeting
Draft every shot at low fidelity first. Rough drafts at reduced resolution and shorter duration let you validate composition and pacing cheaply. Only promote shots that survive the assembly to high fidelity. This one habit typically saves more time than any prompt trick.
A Repeatable End-to-End Workflow
The following sequence is the backbone of a production that scales beyond a single scene.
Step 1 — Write the beat sheet
Twelve to twenty beats, each one sentence, each one a change. Keep it on one page so you can see the whole arc at once.
Step 2 — Build the shot table
One row per shot with the columns described earlier. Write the dramatic function first, then the visual description. If you cannot name the function, delete the row.
Step 3 — Design each shot in words
Draft a one-paragraph visual description per shot as if you were explaining it to a cinematographer. This paragraph is your source of truth; prompts are derived from it, not the other way around.
Step 4 — Generate a look frame
For each new scene or location, generate a single still look frame that establishes palette, lighting, and texture. Approve it before generating any motion. Approving stills is faster and cheaper than approving clips.
Step 5 — Generate low-fidelity drafts
Generate all shots as short, low-resolution drafts against the approved look frame. Assemble them in order the same day, with sound if possible.
Step 6 — Diagnose in the assembly
Watch the rough cut three times. First for story clarity, second for pacing, third for visual continuity. Mark, do not fix. Fixing during a first watch leads to endless small revisions.
Step 7 — Promote, regenerate, or cut
Promote shots that work, regenerate shots that fail for identifiable reasons, and cut shots that never had a dramatic function. Repeat until the sequence reads.
Assembling Prompts That Preserve Directorial Intent
A reliable prompt is structured, not poetic. Use a consistent order so you can debug one variable at a time.
A six-part prompt skeleton
Subject and action; setting and time of day; framing and lens effect; camera movement; lighting and palette; texture and technical notes. Example structure: "A middle-aged woman in a soaked wool coat steps off a curb — narrow alley at dusk after rain — medium shot, compressed background — slow lateral tracking right — cool ambient light with warm sodium practicals — slight handheld drift, filmic grain, shallow depth of field." Each clause maps to one production decision.
Controlling motion without overloading it
Describe one primary movement and, if needed, one secondary micro-movement. If a shot needs three movements, it is three shots. Overloaded motion prompts are the most common cause of unstable output.
Using negative and constraint language sparingly
Negative prompts help with persistent artifacts such as warped hands or text-like smearing, but long lists of exclusions dilute the signal of what you actually want. Keep constraints short and specific to failure modes you have actually observed.
Iterating one variable at a time
When a shot fails, change exactly one clause and regenerate. Changing four clauses at once gives you a better clip with no understanding of why, which means you cannot repeat the success tomorrow.
Continuity and Multi-Shot Coherence
Continuity is the difference between a sequence and a slideshow. Three categories matter most.
Character consistency
Maintain a reference sheet per principal character: front, three-quarter, and profile views, plus two wardrobe variations. Reuse the same reference image across all shots in a scene, and describe the character identically in every prompt — same hair length, same coat, same age range. Drift is cumulative, so a small inconsistency in shot two becomes a different person by shot nine.
Prop and wardrobe continuity
Track objects that change state: a cup that empties, a jacket that comes off, a bandage that appears. Note the state in the shot table so the prompt for each shot reflects the correct moment in time.
Spatial and temporal continuity
Keep a simple floor plan for each location and mark camera positions on it. This prevents the disorienting effect of a room that rearranges itself between cuts. For time of day, keep a scene-level lighting statement and never improvise a change mid-scene unless the story calls for it.
Editing, Pacing, and Sound as Story Carriers
Generation ends where editing begins, and editing is where most of the perceived quality is created.
Cutting on intent, not on frames
Cut when the audience has just understood something or is about to. Cutting a beat too late feels sluggish; too early feels confusing. In AI footage specifically, cut before the artifacts become noticeable — the moment of confidence is often shorter than the clip.
Pace mapping by beat
Assign an approximate duration to each beat in the beat sheet. A confrontation scene might run eighty seconds; a reveal might run four. When the assembly is running long, cut beats rather than trimming every shot, because trimming uniformly flattens rhythm.
The role of sound design
Sound carries continuity that image generation struggles with. A consistent room tone, a recurring musical motif, and clean dialogue levels will make mismatched shots feel like a single scene. Build a simple sound pass early — even a rough one — because pacing judgments are unreliable against silence.
Color grading for cohesion
Use one grade across the sequence. Bring every shot to a common baseline first, then push toward the scene's palette. Grading after generation is also the fastest fix for shots that came from different model families and do not quite match.
Common Mistakes, Fixes, and a Review Checklist
Mistake: prompting before writing
Fix: refuse to open a generation tool until the shot table has dramatic functions filled in for every row.
Mistake: unlimited coverage
Fix: cap coverage per scene at a number you commit to, and delete anything beyond it before rendering.
Mistake: inconsistent scene language
Fix: copy the same lighting and palette clause into every prompt in a scene. Do not paraphrase it.
Mistake: fixing in the first pass
Fix: watch, mark, and only then revise. Batch revisions by category — composition, motion, continuity — rather than by shot.
A short pre-lock checklist
Does every shot have a named function? Does the sequence read with the sound off? Is screen direction consistent? Do reference images match across shots in a scene? Does the cut point land on comprehension rather than on a full clip? Does the grade unify the sequence? If any answer is no, you have a diagnosis, not a mystery.
FAQ: Practical Questions About AI Story and Shot Design
Do I need a full screenplay before generating?
No. A beat sheet plus scene intent notes is enough for short-form work. A screenplay becomes valuable when dialogue and subtext are doing structural work, because those are hard to retrofit later.
How many shots should a one-minute video have?
Typically eight to twenty, depending on genre. Action and montage sit at the high end; dialogue and atmosphere sit at the low end. What matters more than the count is that every shot has a distinct function.
Is image-to-video always better than text-to-video?
No, but it is more controllable. Use image-to-video when composition and continuity matter; use text-to-video when you need complex motion or an environment you cannot easily author as a still.
How do I keep characters consistent across many shots?
Lock a reference sheet, repeat the same descriptive clause word for word, keep wardrobe and lighting constant within a scene, and check consistency every few shots rather than at the end.
What should I do when a shot will not generate correctly?
Split it. Most failures come from asking one shot to do two jobs — establish a location and deliver a performance, or hold composition and perform a complex move. Two simpler shots usually solve what endless retries cannot.
How long should I spend on pre-production versus generation?
A useful rule for new creators is to spend roughly as much time planning shots as rendering them. Planning feels slower because it produces nothing visible, but it is the reason the render stage proceeds in a straight line instead of circles.
Story and shot design are not obstacles between you and the tool. They are the reason the tool produces something worth watching. Write the beats, build the table, design the look, then generate — and judge every result against intention rather than against surprise.

