Start Free Now
Limited Time Offer: Get 50% OFF Starter & Basic Yearly Plans 🎉

AI Shot Planning Workflow: From Script Breakdown to Final Cut

Sep 27, 2026

What AI Shot Planning Actually Solves

For most of film history, the expensive part of production was capture. Film stock, crew days, lighting trucks, locations, and permits all pushed teams to plan obsessively, because every unplanned setup cost real money. Generative video has inverted that equation. Generating a clip is now cheap and fast, while deciding what to generate has become the bottleneck.

Every generated clip is a fork in the road. You choose a model, a seed, a prompt, a reference image, a camera move, a duration, an aspect ratio, a motion strength, and a style anchor. Multiply those choices by eighty shots and you get a combinatorial space no single person can hold in their head. Teams respond by improvising shot by shot, then wonder why the finished piece feels like a slideshow of unrelated images rather than a film.

A shot planning system exists to collapse that space into a repeatable pipeline. Done well, it does four things:

  • Converts a script into an ordered list of shots with a defined narrative purpose for each one.
  • Maps every shot to the technical requirements that will actually produce it, so generation becomes execution rather than experimentation.
  • Tracks continuity state between shots so characters, wardrobe, props, and lighting do not drift across a sequence.
  • Records why each decision was made, so a revision three weeks later does not quietly undo a choice that solved a problem you have since forgotten.

The failure mode without planning is predictable. You generate a gorgeous opening shot, then spend three hours failing to match it in shot seven. Or you prompt each shot in isolation and the protagonist's jacket changes color in every scene. Or you finish the edit and discover that half your footage is unusable at the delivery aspect ratio.

None of those are model problems. They are information architecture problems, and they are solvable long before you open a generation tool.

The Pre-Production Pipeline: From Script to Shot List

Start with scene intent, not shot counts

Before you number a single shot, write one sentence per scene that answers three questions: who wants what, what blocks them, and what changes by the end. This is the spine of the piece. When you are deep in generation at 2 a.m. and a clip looks beautiful but wrong, this sentence is what tells you to discard it.

A useful rhythm for short-form work is one intent sentence per 15 to 25 seconds of finished runtime. Anything denser becomes a treatments document nobody reads; anything sparser leaves you guessing during generation.

Break the script into beats and coverage

A beat is a change in emotional temperature, not a paragraph break. Mark them. Then map coverage: for each beat, what does the audience need to see? A wide establishing geography? A reaction? A detail insert that carries the plot? Coverage planning prevents the most common AI video mistake, which is generating every shot at the same scale and distance because that is what the model does most reliably.

A practical default coverage pattern for a dialogue beat:

  1. Wide or establishing frame to place the space.
  2. Over-the-shoulder or two-shot to establish relationship and blocking.
  3. Single close-up on the character who is driving the beat.
  4. Insert or detail shot to carry information the dialogue cannot.
  5. Reaction or transition frame to hand off to the next beat.

You will not always use all five, and you may use them out of order, but having the pattern in mind keeps your coverage from collapsing into monotony.

Give every shot a machine-readable spec

A shot list written as prose in a note app is not a shot list. Define a schema and use it consistently. A workable minimum:

Field Purpose
Shot ID Stable reference used in filenames and edits
Scene and beat Where the shot sits narratively
Narrative function What the shot must accomplish
Shot size Wide, medium, close, insert
Camera move Static, push, pull, pan, tilt, orbit, handheld
Duration Target length before trimming
Subject state Wardrobe, hair, injuries, props held
Environment Location, time of day, weather
Lighting key Direction, quality, color temperature
Reference assets Character sheet, mood frame, location plate
Model and settings Which engine, which parameters
Status Planned, generated, selected, locked

Whether you keep this in a spreadsheet, a database tool, or a version-controlled text file matters far less than consistency. The payoff arrives when you need to regenerate "every medium close-up of Mara at dusk" and can actually query your own production.

Keep the shot list short enough to finish

Amateur AI projects die from ambition, not from bad tools. If you can generate and review roughly twenty clips in a focused session, a sixty-shot film is three generation sessions plus pickups plus editing. Plan in blocks of ten to fourteen shots and lock each block before starting the next one. Locking means you have selected a final take for every shot in the block and written down why.

Choosing the Right Generative Model Per Shot

No single engine wins every category. Some excel at photoreal people in motion, some at stylized illustration, some at precise camera moves on static subjects, some at long-duration coherence. Treating one tool as universal guarantees mediocre results somewhere in the film.

Build a routing table instead of a favorite. Here is a decision framework that works across most engines:

Shot type Primary priority What to test first
Talking-head close-up Facial stability Lip sync and micro-expression over five seconds
Wide establishing Geography and depth Parallax and horizon stability during any move
Action beat Motion plausibility Limb count, ground contact, motion blur
Product insert Material accuracy Reflection, logo legibility, edge cleanliness
Stylized montage Consistency of style Whether the look survives across ten seeds
Text or sign in frame Legibility Whether letters survive motion

Apply the three-test rule

Before committing an engine to a shot, run three cheap tests. First, the ugliest case: generate the most physically demanding frame your shot contains, not the hero frame. Second, the duration limit: push to your target length plus two seconds and see where coherence collapses. Third, the boring middle: generate the shot at the exact moment it must cut into the next one, because transitions expose failures that isolated clips hide.

Three tests cost minutes. Skipping them costs an afternoon of editing around an unusable take.

Treat routing as policy, not mood

Write your routing decisions into the shot list before generating anything. "Action beats go to the motion-focused engine; product inserts go to the detail-focused engine" is a policy. "I'll try whichever sounds good" is a mood, and moods produce films with three visual dialects that never reconcile.

One more criterion that is easy to overlook: how editable is the output? Some engines return only a finished clip. Others return intermediate data, masks, or depth passes you can composite with. If a shot will need compositing work, choose the engine that gives you something to composite with, even if its raw output looks slightly worse.

Directing Camera Language Without a Camera

Camera language is the fastest way to make generated footage feel intentional rather than accidental. You are not operating a camera, but you are still making the same decisions a director of photography would make, and models respond to them in fairly predictable ways.

Move vocabulary worth using

  • Static. Underrated. Static frames cut together cleanly and let the audience read performance. Use them for emotional beats and for shots that will carry on-screen text.
  • Push in. Increases tension and intimacy. Keep it slow, or the model will overshoot and the frame will drift off-subject.
  • Pull out. Reveals context. Excellent for scene endings because it lands on a wide frame that cuts easily.
  • Pan and tilt. Useful for establishing geography, risky for character work because subjects can smear at the frame edges.
  • Orbit. Impressive but expensive. Orbits distort backgrounds and break facial consistency at roughly the halfway point on most engines.
  • Handheld. Adds realism. Ask for subtle drift rather than aggressive shake, because generated shake tends to look rhythmic and synthetic.

Do the pacing math before generation

Duration is a creative decision, and models reward you for respecting their limits. A three-second clip can hold one idea. A five-second clip can hold an action plus a reaction. An eight-second clip can hold a small arc, but only if the action is simple and the camera is calm.

If your shot needs twelve seconds of continuous action, do not ask for twelve seconds. Ask for two six-second clips that cut on movement, with the character's position and momentum matched across the join. Cutting on action hides the seam and gives the editor freedom later.

Cut on action, not on stillness

When you are planning a sequence, mark the exact frame where each shot hands off to the next. Handoffs on motion, a turn, a step, a door closing, a hand entering frame, read as continuous even when the two clips came from different seeds on different days. Handoffs on stillness expose every inconsistency in lighting and wardrobe.

Continuity Across Shots: The Hardest Problem

Consistency is where most AI video projects visibly fall apart. Audiences forgive a slightly odd hand; they do not forgive a character whose face changes shape between two shots in the same conversation.

Character consistency

Lock a character sheet early: three to five reference frames showing front, three-quarter, and profile views under neutral lighting, plus a wardrobe description written in fixed language you paste into every prompt. Avoid synonyms. If the character wears a "charcoal wool overcoat," never call it a "dark jacket" in a later prompt, because the engine will treat those as different garments.

If the project is long enough, invest in a trained character model or a consistent reference-conditioning workflow. It costs setup time and returns visual stability across dozens of shots.

Prop and wardrobe continuity

Keep a state column in your shot list that records what the character is holding, wearing, and carrying at the start of each scene. Update it whenever the story changes that state. This sounds bureaucratic until the moment you realize shot forty-two must show the same bandaged hand as shot nineteen.

Lighting and environment continuity

Group shots by lighting setup during generation rather than shooting the story in order. Generating all your dusk exteriors in one session keeps the color temperature and shadow direction closer than generating them across three weeks of separate sessions. Generation order and story order do not have to match; only the edit cares about story order.

Build a continuity sheet you actually check

A one-page continuity sheet beats a fifty-page bible every time. Include character appearance, wardrobe by scene, props by scene, time of day, weather, and any injury or damage state. Read it before each generation session and after each edit pass.

Handling Complex and VFX-Heavy Shots

Some shots simply cannot be generated in one pass. Recognising them early saves days.

Decompose into plates and elements

A shot of a character walking past a burning building is three problems: the character, the fire, and the environment. Generate or source them separately, then composite. On most projects a composited shot built from three simpler generations looks better than a single ambitious generation, because each element gets an engine suited to it.

Apply the one-impossible-thing rule

Every generated shot should contain at most one element that the engine is likely to struggle with: a complex hand interaction, a reflective surface with a visible logo, a crowd, an unusual creature. Two hard elements multiply failure rates rather than adding them, and retries get progressively more expensive in time.

Know when not to generate

Generated footage is the wrong tool for a locked-off product beauty shot that a still camera can capture in ten minutes, for legal documents, and for anything requiring exact text. Use stock footage, stills with subtle parallax, or real capture for those. Choosing the right medium per shot is a directing decision, not a concession.

Smoke-test before the full sequence

For any sequence with VFX, generate one rough version of the hardest shot at low quality before committing to the sequence. If the hard shot does not work, restructure the sequence now, while the cost is one clip rather than twenty.

Editing, Sound, and Review Gates

Edit early, generate late

Build a rough cut from placeholder material, even black frames with correct durations and a scratch voiceover. Editing the timing first tells you which shots you actually need, and how long each one must hold. Generating before you have timing means generating shots you will cut, and discovering too late that your pacing beats do not exist.

Sound carries generated footage

Ambience, foley, and music do more continuity work than any visual fix. A room tone bed under an entire scene glues clips with mismatched lighting. Footsteps, cloth movement, and object handling make motion feel physically real even when the render is not.

Naming and versioning discipline

Adopt a filename convention on day one: project, scene, shot, version, status. A shot reviewed and rejected should remain on disk with a status marker, not be deleted, because revisions often resurrect the earlier take. Keep a short decision log per shot: what you tried, what failed, why the selected take won. Future you will not remember, and the log is faster than regenerating.

Use review gates, not endless tinkering

Set three gates. Gate one: story and shot list locked. Gate two: all shots generated and selected. Gate three: picture and sound locked. Between gates, no scope changes. Teams without gates produce projects that are perpetually ninety percent complete, because every new idea restarts the generation phase.

Common Mistakes and How to Avoid Them

Generating before the shot list exists. Fix: write the list first, even a rough one. Decisions made on paper cost nothing.

Using one engine for everything. Fix: build a routing table and respect it.

Prompting with synonyms. Fix: freeze descriptive language for characters, wardrobe, and locations in a shared glossary.

Chasing a shot that will not resolve. Fix: cap retries at a number decided in advance, usually five to seven. If the shot fails at the cap, change the shot, not the engine.

Ignoring duration limits. Fix: plan around each engine's coherent clip length and bridge with cut-on-action edits.

Over-generating. Fix: if a shot exists for coverage and never survives the edit, delete it from the plan.

Neglecting aspect ratio and specs. Fix: confirm delivery resolution, frame rate, and safe areas before generation. Cropping later destroys compositions you carefully planned.

Ignoring continuity until the edit. Fix: use the continuity sheet during generation sessions, not after.

Skipping the audio pass. Fix: budget as much time for sound as for picture polish. It changes perceived quality more than another generation pass will.

No decision log. Fix: thirty seconds of notes per shot prevents hours of re-litigating settled choices.

A Worked Example: A Ninety-Second Product Story

Here is how the workflow looks on a realistic small project: a ninety-second piece about a fictional commuter backpack, intended for a brand page and short-form social cutdowns.

Intent sentences. Four scenes: a commuter leaves before dawn and struggles with a cheap bag; the bag fails, spilling contents on a platform; the character switches to the new bag; a final sequence shows the bag handling rain, a packed train, and a lift to an office.

Coverage plan. Fourteen shots across the four scenes, roughly six seconds each, with two shots held longer for the emotional beats:

  1. Wide, pre-dawn street, push in slowly on the character's back (establishing).
  2. Insert, old bag's zipper straining (foreshadowing failure).
  3. Medium tracking shot, character walking to the platform (rhythm).
  4. Close-up, hand gripping a broken strap (turning point).
  5. Insert, contents scattering on the platform floor (consequence).
  6. Wide, character kneeling to gather items (low point).
  7. Product insert, new bag on a table, static frame, soft key light (introduction).
  8. Detail, water beading on fabric (proof point one).
  9. Medium, character boarding a crowded train, bag against the body (proof point two).
  10. Close-up, laptop sliding into a padded sleeve (proof point three).
  11. Insert, shoulder strap adjusting in one motion (feature highlight).
  12. Wide, office lift doors opening, character stepping out composed (resolution).
  13. Product beauty frame, static, slow pull out (brand moment).
  14. End card frame with on-screen text (delivery).

Routing decisions. Shots 1, 3, 6, 9, and 12 are human-motion driven and go to the engine with the strongest figure realism. Shots 2, 4, 5, 8, and 11 are detail and material shots and go to the engine that handles texture and reflection best. Shots 7, 10, and 13 are product frames and are generated from still photography with subtle parallax, not full generation. Shot 14 is assembled in the edit with motion applied to a still, so the text stays legible.

Continuity state. Character: charcoal overcoat, olive scarf, brown leather messenger strap on the new bag. Old bag: navy nylon, broken left strap from shot 4 onward. Environment: pre-dawn blue for shots 1 through 3, fluorescent platform for 4 through 6, warm interior for 7, 8, 10, 11, 13, daylight for 9 and 12.

Generation order. All pre-dawn shots in one session, all platform shots in one session, all warm interior product shots in one session. Story order is irrelevant here; lighting consistency is not.

Edit assembly. Rough cut with scratch narration at day two, full generation days three and four, sound design on day five. Music and ambience work is what makes the platform sequence feel like one continuous moment rather than six clips.

Social cutdowns. Because each shot was planned with a clear narrative function, the vertical cut can drop shots 2, 6, 10, and 11 and still hold the story in thirty seconds. Teams that never wrote narrative functions discover at this stage that every shot was load-bearing, and the cutdown becomes an argument instead of an edit.

FAQ

How many shots should a first AI video project have?

Aim for twelve to twenty finished shots. That is enough to tell a short story and few enough to finish. Complexity per shot matters more than count: a project with fifteen simple static shots will succeed more often than one with eight ambitious shots requiring compositing.

Do I need a storyboard, or is a shot list enough?

A shot list is the mandatory artifact; storyboards are optional and most useful for action and VFX sequences where spatial relationships matter. For dialogue and product work, a written shot spec plus reference frames usually communicates enough.

How do I keep a character consistent across many shots?

Three practices do most of the work: a fixed reference sheet, a frozen glossary of descriptive language you paste verbatim into every prompt, and generating shots featuring that character in closely grouped sessions rather than weeks apart. Trained character models add more stability if the project justifies the setup.

What is a realistic retry budget per shot?

Decide a cap before you start, usually five to seven attempts. If the shot still fails, the shot is the problem, not the engine. Simplify the action, shorten the duration, change the camera move, or split it into two shots. Chasing a stubborn shot is the single biggest time sink in AI video work.

Should I generate in story order?

No. Generate grouped by lighting, location, and character state for consistency, then assemble in story order during editing. Story order governs the edit; generation order should serve continuity.

How do I handle text and logos in generated video?

Usually do not generate them. Add text, signage, and product logos in the edit or in a compositing pass. Generated lettering warps under motion and is nearly impossible to fix reliably, while clean overlay text is trivial and always legible.

When is a shot better handled without generation at all?

When it needs exact text, precise product geometry, brand-accurate color, or a legally meaningful document in frame. Still photography with parallax, screen capture, or stock footage often beats generation for those cases and integrates invisibly once graded.

How much of my total time should go to sound?

Plan on roughly one fifth to one quarter of the schedule for sound if you want the piece to feel professional. Ambience, foley, and music are the cheapest tools available for making separately generated clips feel like one continuous world.

What does a good shot list actually look like in practice?

It fits on one or two screens, has one row per shot, and contains the technical spec plus the narrative function. If your list requires scrolling through paragraphs of prose to find a camera move, it is documentation rather than a working tool.

How do I keep a larger team from drifting apart?

Single source of truth, one naming convention, and three fixed review gates. Everyone should be able to answer, from the same document, what shot is next, what it must accomplish, and whether it has been approved. When those three questions require a meeting, the workflow has already failed.

Alexander

Alexander