Why AI-assisted direction changes the production pipeline
For most of film history, the screenplay was a private document. It circulated among a small crew, then disappeared into the shooting schedule. The audience never saw it. The shot list was a set of private notes that a director scribbled in the margins and then threw away.
Generative video flipped that relationship. Now the script and the shot description are the input to the camera. A model does not watch a rehearsal and adapt. It reads a description and produces pixels. That means every ambiguity you tolerated in a traditional screenplay — "they argue, tension builds" — becomes a visible failure on screen. The model has no idea what tension looks like unless you describe the body language, the framing, and the light.
This shift has three practical consequences.
First, pre-production becomes production. The quality of your shot descriptions determines the quality of your footage more than any model setting. A beautifully written scene with vague shot notes produces generic footage. A mediocre scene with precise shot notes produces usable footage.
Second, you need a unit of work. Traditional editors work in clips and timelines. AI video work is easier when you think in shot cards: self-contained, versioned descriptions that can be regenerated independently without breaking the rest of the sequence.
Third, continuity becomes a document problem, not a memory problem. A human camera operator remembers what the lead actor was wearing. A model does not. If continuity is not written down in a reusable form, every shot drifts.
The rest of this guide lays out a workflow you can run end to end: breaking down a script, writing director-grade shot descriptions, choosing tools per shot, and reviewing the output without losing your mind.
What a video model actually needs from a shot description
Most disappointing AI video outputs are not model failures. They are specification failures. The model did something reasonable with an underspecified request.
The five anchors of a usable shot description
Every shot description should answer five questions explicitly:
- Subject — who or what is on screen, described physically, not dramatically. "A woman in her fifties, grey braid, oil-stained apron" beats "a tired mother."
- Action — one continuous physical action, present tense. If you need two actions, you probably need two shots.
- Camera — framing, height, angle, and movement. "Low angle, chest-height, slow dolly right, 35mm equivalent" is a specification. "Cinematic" is not.
- Light — source, direction, quality, and color. "Single practical lamp camera-left, hard falloff, warm tungsten, deep shadows on the right side of the face."
- Atmosphere — weather, particles, texture, and environment behavior. Smoke, dust, rain, heat shimmer, steam. These details create motion in otherwise static frames.
If a shot description omits the camera anchor, the model will choose one for you — usually a static medium shot, centered, with flat lighting. That default is the single most recognizable signature of lazy AI video.
Duration, pacing, and the physics problem
Generative models are much better at look than at physics. They handle a slow push-in on a face beautifully. They struggle with an object being handed from one character to another, a door closing at a specific moment, or a character walking a precise distance and stopping.
Practical workaround: write shots that end before the hard physics problem begins. Cut on the reach, not on the handoff. Let the next shot start after the object has changed hands. This is basic coverage discipline and it hides model limitations without drawing attention to them.
Also specify duration and what happens at the end of the shot. "Three seconds, ending as her hand reaches the door handle" gives the model a target. Without an end state, generated clips tend to drift toward a meaningless resting position.
What models cannot infer
Off-screen context is invisible. If your script says "she is hiding something," the model has no access to the previous scene. You must translate subtext into visible behavior: a pause before answering, eyes flicking to the left, hands folded too tightly.
Similarly, models cannot infer style consistency across shots. If you want a sequence to feel like one film, style has to be repeated in every shot description or carried by a reference image. Assume nothing is remembered.
The shot card system: a repeatable unit of work
The most common organizational mistake in AI video production is treating prompts as disposable text in a chat window. You generate something decent, you tweak it, and three days later you cannot reproduce it.
A shot card fixes that.
The six-line shot card template
SHOT: 02B
SCENE: Kitchen, night
SUBJECT: Woman, 50s, grey braid, oil-stained apron
ACTION: Turns from the stove, wipes hands on apron, looks off-frame right
CAMERA: Medium close-up, eye level, slow 10% push in, 40mm
LIGHT: Practical lamp camera-left, hard falloff, warm tungsten
NOTES: Steam from pot drifting through frame; she does not blink
The first two lines identify the shot. The next four are the five anchors compressed into readable form. The notes line carries atmosphere and performance detail.
Two rules make this work. First, one action per card. Second, the card must be readable out loud without losing meaning — if you cannot say it clearly, the model will not render it clearly.
Naming and versioning
Use scene-shot-take naming: S02-B-t03. Keep every take. When a later shot needs to match a character's look, you will want to diff the descriptions and find what changed. Without version history, matching becomes guesswork.
Keep a single spreadsheet or a plain text file per scene. Columns: shot ID, description, tool used, reference image, status, and a short note about what was wrong with rejected takes. That last column is more valuable than it sounds — after twenty shots you will start seeing the same failure pattern repeat, and you can fix it at the description level instead of regenerating blindly.
Scene breakdown: reading a screenplay for directorial intent
Before writing any shot card, break the scene down. Three passes are usually enough.
Pass one: find the turn
Every scene has a moment where something changes — a decision, a revelation, a refusal. Mark it. That moment deserves the tightest coverage: close-up, minimal camera movement, full attention. Everything before it is approach; everything after is consequence.
Pass two: map beats to coverage
Write the beats in plain language, then assign a shot size to each. A useful default progression:
- Establishing — wide, slow movement, environment dominant
- Approach — medium, subject entering frame or settling
- Pressure — medium close-up, camera begins to move or tighten
- Turn — close-up, static or micro-movement only
- Consequence — wide again, or a held close-up after the other character exits
This is not a formula to follow rigidly, but it protects you from the most common AI short film problem: every shot at the same distance with the same neutral energy.
Pass three: mark inserts and cutaways
Inserts are your continuity insurance. A close-up of hands, a clock, a phone screen, a door handle. They are cheap to generate, easy to keep consistent, and they buy you transitions when two consecutive character shots refuse to cut together cleanly.
Plan at least two inserts per scene. In practice they will save a sequence more often than a hero shot will.
Writing prompts that read like director's notes
A prompt is not a spell. It is a set of constraints. The clearer the constraints, the less the model has to invent.
Sentence order carries weight
The subject and action should come first. Models tend to weight early tokens more heavily, and more importantly, this mirrors how a human reads a shot description. Lead with what matters, then qualify with camera and light.
Weak: "Cinematic moody kitchen scene, dramatic lighting, 4k, film grain, woman cooking, sad."
Strong: "A woman in her fifties turns from a stove and wipes her hands on an apron. Medium close-up at eye level, slow push in. One practical lamp on the left, warm tungsten, hard shadow falloff on the right. Steam drifting through the frame."
The second version is longer, but it is also specific, and specificity is what reduces retries.
Negative constraints are cheap and effective
Add a short list of things you do not want: no lens flares, no slow-motion, no smiling, no camera shake, no text overlays, no extra people in frame. Keep it short. Long negative lists tend to confuse models and dilute the positive description.
Reference frames and style anchors
When look matters more than motion, use an image reference. Generate one strong still of your character in the correct wardrobe and lighting, then reuse it. Reference images solve two problems at once: they lock appearance and they reduce the number of words you need for style.
When motion matters more than look, skip references and write a longer, more physical action description. Choose one priority per shot rather than trying to win both.
Continuity across dozens of generated shots
Continuity is where AI video projects live or die. Fortunately, it is a documentation problem, and documentation problems have predictable solutions.
Character sheets
For each recurring character, write a fixed paragraph: age range, build, hair, wardrobe, distinguishing features, and the exact phrasing you will use to describe them. Copy that paragraph into every shot card they appear in. Do not paraphrase it. Paraphrasing is how a character's jacket changes color between shots.
Location bibles
Same idea for places. A location entry should specify architecture, materials, dominant colors, light sources, and time of day. If a scene happens at sunset, write the exact phrase once and reuse it — don't alternate between "golden hour" and "dusk" across shots, because the model will render two different times of day.
Color, grain, and lens continuity
Decide early on a look package: lens character, grain amount, contrast curve, and a three-color palette. Write it once, then append it to every shot card, even when it feels redundant. Redundancy is the point. Models do not carry style forward on their own.
The one exception: when you deliberately want a shot to break continuity — a flashback, a memory, a POV shift — change exactly one variable in the look package and call it out explicitly in the notes. One variable, not three.
Choosing the right tool for each shot
Different shots need different strengths. Rather than committing to one generator for a whole project, route shots based on what they demand.
Use these criteria:
- Motion complexity — simple push-ins and pans are broadly handled well; complex choreography is not. Route complex motion to tools that specialize in camera control, or redesign the shot.
- Human performance — dialogue, subtle facial work, and eye movement have narrow tolerances. Test early with a short clip before committing a whole scene.
- Duration — long continuous takes are harder than short ones. If a tool struggles past four seconds, split the shot rather than fighting it.
- Style fidelity — stylized, illustrated, or painterly looks often hold up better than photoreal faces. Lean into the strength of the tool instead of pushing against it.
- Iteration cost — how fast can you see a result? A tool that produces a rough draft in seconds is often more valuable in pre-production than a slow tool that produces a polished frame.
- Editability — can you extend, outpaint, or re-time the output? Shots you can adjust in post are worth more than finished-looking clips you must regenerate from scratch.
A practical division of labor: use fast, cheap tools for exploration and animatics, mid-tier tools for inserts and atmosphere, and high-fidelity tools only for hero shots where the audience will look closely. Most sequences have two or three hero shots and fifteen supporting shots. Spend accordingly.
Worked example: the opening of a short film
Suppose you are opening a ninety-second short with a scene in a bakery kitchen at night.
Step 1 — Script. One paragraph: a baker finishes a batch, hears something outside, and decides whether to open the back door.
Step 2 — Beats. Finish work. Hear sound. Hesitate. Decide. Cross the room. Open door.
Step 3 — Coverage plan.
- S01: Wide of the kitchen, flour dust in the air, slow dolly right. Establishes space and mood.
- S02: Insert — hands pressing dough, close-up, static.
- S03: Medium close-up of the baker, eye level, static. She stops moving.
- S04: Insert — her head turning slightly, tight framing, shallow focus.
- S05: Wide, static. She stands still, listening. Off-screen sound implied by her stillness.
- S06: Medium, slow push in as she walks toward the back door.
- S07: Close-up of her hand hovering over the latch. Static, held.
- S08: Wide exterior through the door glass. Nothing but darkness and rain.
Step 4 — Shot cards. Write all eight in the six-line format with a shared look package and a single character paragraph copied verbatim into S03, S05, S06, and S07.
Step 5 — Generate in order of risk. Generate the hardest shots first: the character close-ups. If the face does not hold, you want to know before you have built a sequence around it. Inserts are easy and can be made last.
Step 6 — Assemble and trim. Cut to slightly under your intended runtime. AI-generated shots frequently feel slower than they read on the page, so trimming two frames off each cut usually improves rhythm.
Step 7 — Fix at the description level. When a shot fails twice, do not keep regenerating. Rewrite the shot card: simplify the action, change the shot size, or add an insert to cover the gap.
Common mistakes and how to fix them
Every shot is a medium shot. Fix by writing an explicit shot size for every card and checking the distribution across a scene. You want wide, medium, close, and insert represented.
Vague emotional direction. "Sad," "tense," and "determined" are not renderable. Replace with physical behavior: shoulders dropped, jaw tight, one hand gripping a counter edge.
Style drift between shots. Caused by paraphrasing. Fix by locking a look package paragraph and appending it verbatim.
Overlong clips. Models lose coherence as duration increases. Fix by splitting the shot into two cards and cutting between them.
Too many characters in frame. Interaction between multiple characters is the weakest area for most generators. Fix with single coverage, reaction shots, and over-the-shoulder framing — the same solution live-action directors use on tight schedules.
No reference images for recurring characters. Fix by generating a character sheet image first and reusing it as a reference where the tool supports it.
Chasing a perfect single take. Fix by accepting ninety percent and covering the remaining ten percent with an insert or a cut.
Quality control, iteration, and FAQ
Build a short review checklist. Watch each generated shot three times: once for look, once for motion, once with the sound off to confirm the action reads without dialogue. Reject on any of three grounds — wrong subject, wrong motion, wrong light. Do not reject because "it feels different," unless you can name the difference.
Iterate in a fixed order. Change one variable at a time: action, then camera, then light, then style. Changing three things at once teaches you nothing about which one was wrong.
Keep a failure log. Two lines per rejected take: what you asked for and what you got instead. After a scene or two, the pattern becomes obvious and your first-pass hit rate climbs noticeably.
Frequently asked questions
How many shots should I plan per minute of finished video? For fast-paced narrative work, six to twelve shots per minute is a reasonable starting range. Slower, atmospheric pieces may use four to six. Plan more inserts than you think you need; unused inserts cost almost nothing.
Should I write dialogue in shot cards? Keep dialogue in the script, not the shot card. Shot cards describe what the camera sees. If a tool supports lip-sync, you will feed dialogue separately and keep the shot description focused on framing and performance.
How do I handle a character who appears in every scene? Generate one strong still that defines their look, save the exact descriptive paragraph, and treat both as canon. Every shot card that includes them should copy that paragraph without edits.
What if a tool renders beautiful frames but I cannot control camera movement? Lean into static shots and use editing to create rhythm. Cut frequency, insert placement, and music timing do more for perceived motion than a camera move does.
Is it better to write the whole shot list before generating anything? Write the coverage plan for the scene you are working on, not the whole film. Generate, learn what the tools handle well, and adjust the plan for the next scene. Shot lists improve dramatically after the first hour of real generation.
How do I keep a sequence from feeling generic? Specificity in three places: wardrobe with visible wear, lighting with a stated source and direction, and one small physical behavior per shot that a model would not default to. Those three details do more for personality than any style keyword.


