Start Free Now
Limited Time Offer: Get 50% OFF Starter & Basic Yearly Plans 🎉

AI Story Planning and Shot Design: A Practical Workflow

Oct 6, 2026

Why pre-production decides whether your AI video works

Most disappointing AI videos fail long before the first frame is rendered. The failure happens in the planning stage: a vague idea, no shot list, no sense of rhythm, and prompts written one at a time with no memory of what came before. Generative video models are now good enough that the bottleneck has moved. The question is no longer whether a model can render a convincing image of a person walking through rain. The question is whether you know what you are building, in what order, and how each piece connects to the next.

Traditional film pre-production exists to de-risk a shoot. You break down the script, scout locations, decide coverage, and rehearse so that expensive shooting days are not wasted on guesswork. AI production inverts part of that equation: rendering is cheap and fast, but iteration carries a different cost. You pay in re-rolls, in the slow drift of a character's face across dozens of clips, and in the mental overhead of tracking which generation was the good one. A disciplined plan converts that chaos into a repeatable pipeline.

A useful mental model: an AI director layer sits between your intent and the render engine. It listens to your logline, asks structural questions, proposes a beat sheet, then translates beats into shots with defined framing, movement, and duration. That translation step is where quality is won or lost. Skip it and you end up with a folder of beautiful clips that refuse to cut together.

What an AI director layer actually does

An AI director is not a single feature. It is a set of planning capabilities that wrap around video generation. Understanding the layers helps you decide where to lean on automation and where your own judgment still matters most.

Narrative structure guidance

The first layer turns an idea into a shape. You give it a logline, a genre, a target runtime, and a tone. It returns a beat sheet: an opening image, a catalyst, a midpoint reversal, a climax, a resolution. Different frameworks suit different formats. A three-act structure works for a two-minute brand film. Kishotenketsu, the four-part introduction-development-twist-conclusion pattern, often fits short-form vertical video better because it does not depend on conflict. Save-the-cat style beat sheets are useful when you need emotional escalation in under ninety seconds.

The practical value here is not creative genius. It is coverage of the boring questions. Who wants what? What stands in the way? What changes between the first scene and the last? Answering those early prevents the classic AI video failure mode: gorgeous footage of a character doing something vaguely interesting for eighty seconds with no reason to keep watching.

Shot design and scene simulation

Once the beat sheet exists, the next layer converts beats into shots. This is shot design: the deliberate choice of framing, lens, camera movement, subject blocking, and shot duration for every beat of the story. A well-built shot list tells a generator exactly what to produce and tells an editor exactly how the pieces interlock.

Simulation is the underrated half of this. Before spending time on final renders, you can build a rough animatic from still frames, time each shot to your target duration, and watch the sequence play. Roughly half of all editing problems are visible at this stage: a shot that runs two seconds too long, a jump cut that breaks screen direction, an establishing shot placed after the audience already figured out where they are. Fixing those in a storyboard costs minutes. Fixing them after thirty generations costs an afternoon.

Character consistency across shots

Consistency is the hardest technical problem in AI video, and planning is the cheapest place to solve it. A character bible captures the details that must never drift: face shape, hair, age, wardrobe, signature props, posture, vocal tone. Reference sheets — a front view, a three-quarter view, a profile, and two or three emotional expressions — give image and video models something stable to anchor to.

Techniques that help include reusing a fixed seed across a scene, conditioning on a locked reference image, training a lightweight identity adapter on a small set of approved images, and blending multiple references when one image alone produces a generic face. Wardrobe and color locks matter just as much as the face. If your protagonist wears a red jacket in shot four, it needs to be the same red in shot nineteen.

Building a shot list that a video model can follow

A shot list is only useful if it is specific. Vague entries like "wide shot of city" produce vague results. Each row should carry enough information that a stranger could generate the shot without asking you a single question.

Coverage patterns worth pre-planning

Professional coverage follows patterns because patterns cut well. An establishing wide shot orients the audience. A master shot holds the whole scene in one frame. Medium shots carry dialogue and body language. Close-ups carry emotion. Inserts — hands, objects, screens, a door handle turning — provide rhythmic punctuation and give editors something to cut to when continuity breaks.

Two rules deserve special attention in AI work. The first is the 180-degree rule: keep the camera on one side of the action line so characters maintain consistent screen direction. When a model generates a shot from the opposite side, your conversation scene suddenly looks like two people talking to the same wall. The second is eyeline matching. If a character looks slightly left of frame in a close-up, the reverse shot must show the other character looking slightly right.

A notation system for lens, movement, and duration

Consistency in notation pays off when you hand prompts to different models. A compact format works well:

Field Example Why it matters
Slug SCENE 04 — ROOFTOP Groups shots by location and setup
Shot size Medium close-up Drives prompt vocabulary directly
Lens 50mm, shallow depth Controls compression and background blur
Movement Slow dolly in Prevents random model motion
Duration 6 seconds Sets generation length and edit rhythm
Action She turns, sees the light Gives the model a clear verb
Continuity Red jacket, wet hair Lists details that cannot drift
Audio Low synth drone Reminds you to plan sound early

Shot duration deserves real thought. AI models often default to slow, drifting motion, which reads as lethargic when every shot runs eight seconds. Cutting most shots between three and five seconds, reserving longer holds for emotional beats, instantly makes AI footage feel more intentional.

Choosing the right video model for each shot

Model choice is a creative decision, not a technical afterthought. Different engines have different temperaments. Some excel at photoreal human faces and skin texture. Some handle stylized, painterly motion better. Some offer precise camera control. Some hold coherence across longer durations, while others produce a spectacular three seconds and then dissolve.

Match model temperament to shot intent

Build a small personal test reel. Take the same prompt — a person walking through a doorway into warm light — and run it through four or five different engines. Watch for three things: how the face holds, how the camera behaves, and how the motion resolves at the end of the clip. Save the results with clear labels.

Then match strengths to shot types. Close-ups of faces belong with the engine that renders skin and eyes best. Establishing shots and landscapes can go to the engine with the richest detail and widest dynamic range. Fast action and stylized sequences can go to whichever engine handles motion blur and physical plausibility without melting. This kind of routing turns a single pipeline into a small studio with specialists.

Frame control and keyframe conditioning

Frame control is the bridge between shot design and execution. Keyframe conditioning — supplying a first frame and sometimes a last frame — gives you precise control over where a shot begins and ends. Image-to-video generation locks the opening composition. Motion guidance tools let you sketch a camera path or specify that the camera pushes in, pulls out, or arcs around a subject.

Use these deliberately. A shot that must cut seamlessly into the next one benefits enormously from a locked last frame that matches the following shot's first frame. A shot that must reveal something benefits from a two-keyframe approach: the closed door, then the open door. When you plan keyframes as part of the shot list, generation becomes assembly rather than gambling.

A step-by-step workflow from logline to first assembly

  1. Write the logline and the ending first. One sentence for the premise, one sentence for how it resolves. If you cannot state the ending, no amount of shot design will save the piece.
  2. Choose a structure and a runtime. Ninety seconds for social, three minutes for a product story, longer for narrative shorts. Structure choices flow from runtime.
  3. Generate a beat sheet. Eight to fourteen beats for a short piece. Give each beat a one-line description and an emotional target.
  4. Write the character and world bible. Faces, wardrobe, props, locations, palette, and time of day. Add reference images before generating anything.
  5. Convert beats into a shot list. Two to five shots per beat is typical. Assign shot size, lens, movement, duration, and continuity notes.
  6. Build an animatic. Stills in a timeline with the intended durations. Add temporary music. Watch it twice and cut anything that drags.
  7. Lock the script and the shot list. Generation should not be interrupted by story changes. If the story changes, return to step five.
  8. Generate in scene order. Do one location at a time so lighting and wardrobe stay coherent, and keep a running log of seeds, prompts, and reference images.
  9. Assemble, then patch. Cut the rough sequence, mark weak shots, and regenerate only those. Repairing five shots is fast; re-rolling a whole scene is not.
  10. Finish sound and grade. Dialogue, ambience, music, and a unifying color pass do more for perceived quality than another round of generations.

Prompt and keyframe templates that survive iteration

Prompts fail when they change shape between shots. A consistent prompt scaffold keeps your look stable and makes debugging possible, because you know which variable you changed.

A reliable scaffold has seven slots:

  • Subject: who or what, with identifying details from the bible.
  • Action: one clear verb phrase. Two actions confuse the model.
  • Setting: location, time of day, weather, background elements.
  • Lighting: key light direction, quality, color temperature.
  • Optics: lens length, depth of field, framing.
  • Mood and palette: three color words maximum.
  • Motion: how the camera and subject move, and at what speed.

An example: "Woman in her thirties, wet dark hair, red rain jacket, walking toward camera through a narrow alley at night, neon signage behind her, hard key light from the left with cyan rim light, 35mm lens, shallow depth of field, tense and cinematic, teal and magenta palette, slow steady push-in."

Keep every other slot identical when you debug. If the face drifts, change only the identity reference, not the lighting. If the mood is wrong, change only the palette words. One variable at a time is the difference between iteration and noise.

Common mistakes and how to fix them

Generating before the script is locked. The most expensive mistake. Every story change invalidates shots you already produced. Fix: force yourself to write the beat sheet and shot list to completion before opening any generation tool.

Identity drift across a scene. A character's face subtly changes between shots. Fix: lock a reference image, reuse the seed, keep wardrobe descriptions verbatim, and generate a scene in a single session rather than across days.

Style drift across the film. Each shot looks like it came from a different movie. Fix: define a palette of three to five colors, a consistent lighting scheme, and a lens family. Apply a unifying grade at the end as insurance.

Broken screen direction. Characters look like they are addressing the wrong side of the frame. Fix: draw a simple overhead diagram before generating conversation scenes, and mark the action line in your shot list.

Overlong shots. Every clip runs the maximum duration, so the edit feels slow. Fix: set a target duration in the shot list, with a hard rule that only climactic beats exceed six seconds.

Constant camera movement. The model drifts and pans in every generation, which reads as amateur. Fix: explicitly state "static camera" or "locked-off shot" for a large share of your coverage. Stillness creates contrast for the moments that move.

No naming convention. Files multiply and nobody knows which take is approved. Fix: name files by scene, shot, and version — scene04_shot02_v3 — and keep a simple spreadsheet log of prompts and seeds.

A consistency QA checklist before you export

Run this pass on the assembled timeline before final delivery. It catches the majority of issues that audiences notice subconsciously.

  • Faces: identity, age, and hairstyle match across every appearance.
  • Wardrobe: colors and garments are identical shot to shot, including accessories.
  • Props: objects stay in the same hand and the same state of wear.
  • Geography: entrances, windows, and furniture stay on the same side of the room.
  • Screen direction: movement and eyelines remain consistent within each scene.
  • Lighting: the light source stays plausible relative to time of day and location.
  • Palette: no shot introduces an off-brand color that breaks the grade.
  • Motion: camera style stays within the rules you established for the project.
  • Sound: ambience changes match location changes, and music supports rather than overpowers.
  • Pacing: no shot outstays its welcome, and cuts land on the beat.

Frequently asked questions

Do I need a full script before generating anything?
A complete screenplay is not required, but a locked beat sheet is. Your beat sheet defines how many shots you need and what each must accomplish. Without it, you will generate clips that look good individually and cannot be assembled.

How many shots should a one-minute AI video contain?
Twelve to twenty shots is a comfortable range, which averages three to five seconds per shot. Fewer shots feel slow and floaty; many more feel frantic unless the subject is action-driven.

How do I keep a character consistent without training a custom model?
Use a locked reference image, reuse the same seed, describe the character with identical wording every time, and generate an entire scene in one session. Custom identity adapters help at scale, but disciplined reference reuse carries most projects.

Should I use one video engine or several?
Several, if you can manage the workflow. Routing close-ups, establishing shots, and stylized sequences to the engines that handle them best improves results noticeably. Keep a short test reel so routing decisions stay based on evidence.

Why does my footage look amateur even though the images are beautiful?
Usually pacing, screen direction, or camera movement. Beautiful frames that drift, overstay, or contradict each other read as amateur. Tighten durations, lock the camera more often, and enforce continuity.

Can I plan shots entirely in text?
Text gets you about seventy percent of the way. The remaining thirty percent comes from seeing the sequence. Storyboard frames, even rough ones, reveal rhythm problems that no amount of written detail exposes.

How much does shot planning actually save?
In practice, teams that lock a shot list before generating report far fewer wasted generations per finished shot, because each generation has a defined success condition. The saving is not in the rendering, it is in knowing when to stop.

Where to take this next

Treat shot design as the core craft skill of AI video, not a preliminary chore. Start small: one scene, five shots, one locked character reference, one animatic. Study how those five shots cut together, then expand the same discipline to a full piece.

The workflow scales predictably. Logline, beat sheet, character bible, shot list, animatic, generation in scene order, assembly, sound, grade. Every step you skip comes back later as rework. Every step you complete makes the next one faster, until the pipeline runs on decisions instead of luck — and your footage finally looks like it was directed by someone who knew what they were making.

Alexander

Alexander