Start With the Story, Not the Generator
Generating a single impressive shot has become almost trivial. Generating a sequence that holds together for thirty seconds, two minutes, or a full episode is still hard, and the difficulty has very little to do with the generator itself. The hard part is pre-production: deciding which shots must exist, what each one contains, how they cut together, and what must stay identical from frame one to frame last.
That is why an AI-assisted storyboard phase matters more than ever. When a shot takes seconds to produce, the real cost shifts to the decision-making upstream. A weak shot list multiplies into a weak edit. A vague prompt multiplies into a reshoot you cannot schedule. Storyboarding is simply the cheapest place in the pipeline to be wrong.
This guide walks through a complete, tool-agnostic workflow: turning a script into visual beats, writing a shot list a generator can actually follow, translating shots into prompts, locking character and scene consistency, specifying camera language, choosing the right generation mode per shot, and cleaning up the mess that always appears around version seven. Nothing here depends on one specific platform, so you can apply it whether you work with Midjourney, Runway, Kling, Sora, Luma, Pika, Stable Diffusion, ComfyUI, or a mix of all of them.
Build a Shot List the Generator Can Follow
From script beats to visual beats
Start by reading the script out loud and marking every place where the audience learns something new, feels a shift, or changes location or time. Those are your beats. A ninety-second script usually contains six to twelve beats; a three-minute piece rarely needs more than twenty.
Each beat becomes one to three shots. Resist the temptation to write twenty shots for a twenty-second scene. AI video rewards clarity. If you cannot describe a shot in one sentence without using the word "and" twice, it is probably two shots.
What every shot row should contain
Use a spreadsheet or a table. One row per shot, with these columns:
- Shot ID — sequential, stable, never reused after deletion.
- Beat — which story moment it serves.
- Description — one sentence: subject, action, setting.
- Framing — wide, medium, close-up, extreme close-up, insert.
- Camera — static, pan, tilt, dolly, handheld, crane, orbit.
- Duration — target seconds, plus an acceptable range.
- Audio — dialogue, effect, music cue, or silence.
- Continuity notes — wardrobe, props, time of day, screen direction.
The continuity column is the one most creators skip, and it is the one that saves the most time later. If a character holds a red mug in shot three, that mug must exist in shot four, even if it sits on the table out of focus. Write it down before you generate, because you will not remember at 1 a.m. when shot four is finally rendering.
Write Prompts That Read Like Director's Notes
The six-part prompt formula
A reliable AI video prompt is not a paragraph of adjectives. It is a structured brief. Six parts, in a fixed order, will beat a beautiful free-form sentence almost every time:
- Subject — who or what, with identifying details that must persist.
- Action — the single visible verb, in present tense.
- Setting — location, time of day, weather, background activity.
- Framing and camera — shot size plus movement.
- Lighting — source, direction, quality, contrast.
- Look — film stock or grade reference, lens character, grain, aspect ratio.
An example: "A woman in a charcoal wool coat, early thirties, dark hair tied back, walking toward camera along a wet night street. Medium tracking shot, camera pulls backward at walking pace. Sodium streetlights behind her, rim light on shoulders, shallow depth of field. Anamorphic look, cool shadows, warm highlights, subtle grain, 2.39:1."
Notice that nothing in that prompt is a mood word like "cinematic" or "epic." Those words do work, but they do it vaguely. Concrete nouns and camera terms give the model fewer ways to guess wrong.
Describe what you want to see, not what you want to avoid
Negative prompts help, but they are a weak substitute for specificity. If hands keep appearing wrong, do not just write "no deformed hands" — reframe the shot so hands are not the subject, or crop them out entirely. If the model keeps adding a crowd to your empty street, specify the emptiness positively: "deserted street, no pedestrians, closed shutters, parked car with wet roof."
The practical rule: fix composition problems with composition, not with prohibitions. A negation asks the model to subtract from something it has already imagined. A reframe asks it to imagine the right thing from the start.
Keep a style block and reuse it verbatim
Write your look paragraph once, save it, and paste it into every prompt in the sequence. Consistency across shots comes more from repeated phrasing than from any single setting. Change one variable at a time: first the action, then the framing, then the lighting. When you change four things at once, you lose track of which change caused the improvement.
Lock Consistency Before You Scale
Consistency failures rarely appear in the first shot. They appear in the fifth, when the character's face has drifted, the jacket has changed color, and the location no longer matches the establishing shot.
Character sheets
Create a reference sheet for every recurring character before generating any scene. A useful sheet contains:
- One neutral expression, front-facing, evenly lit.
- One three-quarter view.
- One profile.
- One full-body shot showing silhouette and proportions.
- A short written description of age, build, hair, and default wardrobe.
Use these images as references in every shot where the character appears, not just the close-ups. Reference strength should be high for faces and moderate for wardrobe, so the character can react naturally to lighting instead of looking pasted into the scene.
Locations, props, and wardrobe
Do the same for sets. Generate a clean establishing frame of each location in daylight and at night, then reuse those frames as references. Keep a prop list with a one-line description and a photo-like reference for anything the audience will notice twice: a phone, a ring, a car, a specific chair.
Wardrobe deserves special attention because it is the fastest way for an audience to notice a continuity break. If a character changes a jacket between scenes, that should be a story decision, not a generation accident.
Specify Camera Language Explicitly
Camera language is where most AI sequences feel amateur. Model defaults tend toward slow, drifting, slightly floating motion. That is fine for a mood piece and terrible for a dialogue scene.
Movement vocabulary
Use concrete terms and expect them to be interpreted literally:
- Static / locked-off — no movement. Underused and extremely useful.
- Pan — horizontal rotation from a fixed position.
- Tilt — vertical rotation from a fixed position.
- Dolly in / out — camera physically moves toward or away from the subject.
- Truck / track — camera moves laterally, parallel to the subject.
- Crane / boom — camera rises or falls.
- Orbit — camera arcs around the subject.
- Handheld — subtle instability, breathing motion, micro-corrections.
Add speed: slow, moderate, fast. Add motivation: following the subject, revealing the room, settling on the object. Motivation is what separates a camera move that reads as intentional from one that reads as noise.
Pacing and shot duration
Plan durations as ranges, not exact numbers. A close-up of a reaction can hold for three to five seconds. An establishing wide often wants only two. If your edit feels sluggish, the fix is usually shorter shots, not faster motion inside shots.
A useful discipline: write the edit in your head before generating. If a cut is meant to land on a beat, generate the shot so that the interesting action happens in the final third, giving you room to trim.
Choose the Right Generation Mode Per Shot
Not every shot should be made the same way. Matching mode to shot type saves enormous time.
Text-to-video
Best for establishing shots, landscapes, abstract transitions, and anything where no specific identity must persist. Fast, forgiving, and cheap to explore.
Image-to-video
Best for character shots, product shots, and anything requiring a precise composition. Generate or select a still first, approve it, then animate it. This two-step approach gives you a checkpoint: if the still is wrong, you find out before spending time on motion.
Video-to-video and motion transfer
Best for matching a specific performance or move. Shoot a rough reference on a phone, then transfer its motion and timing onto a generated character. This is the most reliable way to get natural body language when the model's default movement feels stiff.
Keyframes and interpolation
When a shot must start and end in exact positions — a product turning from front to back, a door closing — define both ends and let the model interpolate. Interpolation shots are easy to control and hard to make dynamic, so use them for precision, not drama.
| Shot type | Recommended mode | Why |
|---|---|---|
| Establishing wide | Text-to-video | No identity to preserve |
| Character close-up | Image-to-video | Locks face and framing |
| Dialogue exchange | Image-to-video, then edit | Control over eyelines |
| Action insert | Video-to-video | Natural motion |
| Product rotation | Keyframe interpolation | Exact start and end |
A Worked Example: A Thirty-Second Product Film
Here is how the workflow compresses in practice. Suppose you are making a thirty-second film for a stainless steel water bottle.
Beats: quiet morning kitchen, the bottle on the counter; a hand picks it up; the bottle travels in a bag through a city; it is opened on a rooftop at sunset.
Shot list:
- Wide, kitchen counter, morning light, bottle centered. Static. Two seconds.
- Close-up, condensation on steel. Slow dolly in. Two seconds.
- Medium, hand enters frame and lifts the bottle. Handheld. Two seconds.
- Insert, bag zipper closing over the bottle. Static. One and a half seconds.
- Wide, city street, walking with the bag. Tracking shot. Three seconds.
- Medium, rooftop, back to camera, city skyline. Static. Two seconds.
- Close-up, cap twisting open, steam of cold air. Slow orbit. Two seconds.
- Wide, subject drinks, sun flares behind. Static. Three seconds.
- Insert, bottle resting on the ledge, sun setting. Slow dolly out. Two and a half seconds.
The bottle itself is the only element requiring strict consistency, so generate one hero reference image and reuse it in shots 1, 2, 4, 7, and 9. The person's face is mostly hidden or in silhouette, which means you do not need a character sheet at all. That single decision removes most of the consistency risk from the project.
Learning to notice these shortcuts is the real skill. Ask of every shot: what does the audience actually need to identify here?
Common Mistakes and How to Fix Them
Generating before writing the shot list. You end up with forty clips and no sequence. Fix: write all rows first, then generate in order of story importance.
Changing multiple prompt variables at once. You cannot tell what helped. Fix: one variable per iteration, and log what you changed.
Ignoring screen direction. Two characters walking in opposite directions should stay on consistent sides of the frame across cuts. Fix: add screen direction to your continuity column.
Overusing camera movement. Everything swinging and drifting reads as unstable rather than dynamic. Fix: make at least a third of your shots locked-off.
Aspect ratio drift. Some tools quietly output a different ratio than requested. Fix: verify dimensions on every batch and crop deliberately, not accidentally.
Continuity checked only at the end. Fixing a face after twenty shots are generated means regenerating everything around it. Fix: review each beat as a group, not each shot individually.
Treating short clips as finished. Two-second generations rarely contain a usable performance start to finish. Fix: generate longer than you need and trim to the best window.
No naming convention. Files named "final_v3_use_this_one" end projects. Fix: project_scene_shot_take with zero-padded numbers.
Review Loops and Asset Hygiene
Review in beats. Watch shots one through four together, muted, then listen to the audio alone. Problems that are invisible in a single clip become obvious in a sequence.
Keep three tiers of assets: references (character sheets, location plates), working generations (everything you tried), and selects (approved shots with locked versions). Never delete references. Rename nothing after an edit begins, because your edit timeline will break silently and you will not know why.
Set a regeneration budget per shot in advance. If a shot has taken eight attempts, the problem is not the prompt settings; the problem is the concept. Simplify the shot, or cut it and solve the story problem another way.
FAQ
Do I need actual drawing skills to storyboard with AI?
No. A written shot list plus reference images is enough. What you need is the ability to describe framing and movement in words, which is a learnable vocabulary rather than an artistic talent.
How many shots should a one-minute video have?
Usually between twelve and twenty-five, depending on pace. Fast montages can run higher; dialogue scenes run lower.
Which comes first, the script or the storyboard?
Always the script, or at least a beat outline. Storyboarding a vague idea produces beautiful shots that cannot be edited into a story.
How do I stop character faces from changing between shots?
Use a character sheet as a reference image in image-to-video mode, keep the written description identical, and avoid extreme angles where the model has to invent facial structure it has never seen.
Is a locked-off shot ever the right choice?
Frequently. Static shots give the audience a stable frame to read, and they make the moving shots around them feel more deliberate.
What is the fastest way to test an idea?
Generate one still per shot at low cost, arrange them in order, and watch them as a slideshow with rough timing. If the sequence does not work as stills, motion will not save it.
How do I handle dialogue with generated characters?
Generate the performance with clear mouth movement, choose takes where the delivery is readable, and treat the audio as a separate layer you control. Keeping dialogue shots short makes lip-sync issues far less noticeable.
The overall pattern is simple: decide more, generate less, and review in groups. Storyboards and shot lists are not paperwork around the creative work — with AI video generation, they are the creative work. Everything downstream is execution.



