Why planning is the real bottleneck in AI video production
Most unsuccessful AI video projects do not fail because the generator is weak. They fail in the ten minutes before generation, when someone types a loose sentence into a text box and hopes the result will cut together with whatever came before it. The model returns something beautiful, and it is unusable: mismatched eye lines, a jacket that changes color between shots, a camera that drifts when the scene needed stillness, a face that belongs to a different person each time it appears.
A director-style planning layer solves a different problem than a generator does. It reads your script, decides what each shot is for, proposes framing and camera behavior, and stores continuity notes in one place so details stay stable across a whole sequence. Generators are memoryless by design. Planning is what supplies memory, intent, and the discipline to keep both.
This guide lays out a complete pre-production workflow that works with any combination of image tools, text-to-video systems, image-to-video systems, and editors. It is written for two audiences. Beginners get a repeatable process that prevents the most common beginner disasters. Experienced creators get decision criteria and a triage ladder that reduce wasted generations and make reviews faster.
The underlying principle is simple: every shot you design should be cheap to change on paper and expensive to change in pixels. You want to argue about framing while it costs nothing, not after you have spent an afternoon rendering versions that will never be used. Once you accept that principle, the rest of the workflow is mostly bookkeeping, and bookkeeping is what separates a finished film from a folder of attractive clips.
The four artifacts to finish before you generate anything
A tidy AI pipeline produces four documents before a single polished clip exists. If these four stay current, the production stage becomes mechanical rather than emotional.
1. The shot intent list. One line per shot describing who is in frame, what happens, why the shot exists, and roughly how long it lasts. No camera jargon needed yet. This list is your spine: everything else is derived from it.
2. The storyboard frames. Rough still images, one per shot, at one fixed aspect ratio. These are communication instruments, not portfolio pieces. Their job is to expose problems in the sequence while problems are still cheap.
3. The motion and audio notes. Camera movement, subject movement, ambient sound, dialogue beats, and voiceover cues attached to each shot. Keep them on the same row as the shot so you never have to guess which note belongs where.
4. The continuity sheet. Character descriptions, wardrobe, props, locations, time of day, and light direction, written once and then copied word for word into every prompt that needs them.
Add a fifth private artifact that experienced creators keep but rarely show: the risk list. Two or three shots per project will be technically difficult — two people interacting, a moving camera plus dialogue, a complex prop, a crowded street. Name them early, test them first at low resolution, and give yourself time to redesign if they refuse to cooperate.
The loop is iterative. You draft shots, sketch boards, notice a missing beat, add a shot, and re-sketch. Beginners try to skip straight to generation because planning feels like delay. In practice, planning is the fastest part of the project, and it is the only part where a change takes seconds.
From script to shot intents: a repeatable breakdown method
Start by reading your script out loud and marking beats. A beat is any moment where information, emotion, or location changes. Each beat usually becomes one to three shots. If a beat produces six shots, you are probably over-covering; if it produces none, the audience will feel the jump.
Every shot must have a function you can name. The common functions are: establish a place, introduce a character, escalate tension, reveal information, deliver a reaction, transition between scenes, and land a final image. If you cannot state a shot's function in a few words, cut it. Shots without functions are the reason short pieces feel long.
Write each intent as a single sentence with five components: subject, location, action, emotion, and camera intention. Examples:
- Maya stands alone at the empty bus stop, shoulders tight, wide static frame, cold morning light.
- Close on her thumb hovering over an unsent message, handheld, shallow focus, restless.
- The bus arrives and blocks her from view, medium shot, slow lateral camera, resigned.
That sentence is not the prompt you will paste into a generator later. It is a contract with yourself. The prompt gets assembled afterwards from the intent line plus the continuity sheet, which is exactly what keeps fifty prompts feeling like they belong to one film.
A few practical rules make this step faster:
- Design for short clips. Most generated footage reads best between three and eight seconds. Complex motion, multiple characters, and synchronized dialogue tend to degrade past ten seconds. Short by design means easy to assemble later.
- Estimate duration before you write the prompt. Guessing duration during editing is how you end up with a two-minute cut of nine-second holds.
- Add a priority column. Mark each shot must-have, nice-to-have, or expendable. Every sequence has a handful of shots that carry the story and many that decorate it. When time runs short, priority tells you what to cut without destroying the point of the piece.
- Group by location and lighting setup. Generate all the kitchen shots together, then all the street shots. Batching keeps reference frames and prompt language fresh in your head and reduces small inconsistencies.
For sizing: a sixty-second social piece usually needs twenty-five to forty shots. A three-minute narrative short typically lands between eighty and one hundred twenty. Much less than that and the piece feels static; much more and your individual shots are probably too short to read.
Storyboard frames: rough beats pretty
An AI storyboard has one job: answer three questions quickly. Does the shot read? Does the sequence flow? Does the composition support the emotion you want?
Beauty is not on that list, and this is the single most common beginner error. Heavily rendered storyboard frames create emotional attachment, and you will then fight to preserve a gorgeous image that does not serve the edit. Grey blocks, silhouettes, and simple tonal studies communicate just as well and can be revised in seconds instead of minutes.
Use a fixed template for every frame so comparisons are meaningful:
[shot type] of [character description] in [location], [lighting], [color palette], [aspect ratio]
Feed the same character description every single time. Where a tool accepts reference images, attach your character sheet and location plate so the frames stay visually related to each other. Where it does not, keep the wording identical and accept minor variation, because the board only needs to be accurate, not photoreal.
Three more rules that protect the edit later:
One aspect ratio per board. If you need both a vertical cut and a widescreen cut, generate two boards. Cropping a widescreen frame into vertical almost always destroys composition, because the composition was built for the wider shape.
Draw the frame, not the subject. Composition is the actual work of directing: negative space, eye line, horizon placement, and where the viewer looks first. A board that shows only a face centered in an empty square has made no decisions at all.
Board movement, not just moments. For any moving shot, sketch a rough arrow or a second frame showing where the camera ends up. A push-in that starts too tight has nowhere to travel, and you only discover that when the render comes back.
When the board is done, review it as a sequence, not as individual images. Scroll through it fast, at something close to the pace of the final cut. Sequences that collapse at that speed will collapse in the edit too, and it is far easier to add a shot to a board than to shoot a new one.
Continuity: locking characters, locations, and light
Consistency is the biggest technical frustration in AI video, and the fix is bureaucratic rather than magical: write descriptions once and reuse them exactly. Models do not understand intent, but they do respond predictably to identical wording, which means discipline outperforms cleverness here.
The character sheet records age range, build, hair, facial hair, eye color, skin tone, wardrobe with materials, distinguishing marks, and posture. Add reference images from several angles. If a tool supports identity or style adapters, train them from that sheet rather than from random stills pulled off the internet.
The location plate records architecture, furniture, dominant colors, practical light sources, weather, and time of day. One wide reference frame reused as the first frame of every shot in that location improves coherence enormously, because the model is matching a known image instead of inventing a new room.
Then practice language discipline. Never paraphrase a character description between prompts. Copy and paste it. Small wording changes produce large visual changes: describe a jacket as cropped, and you may get one; describe it as fitted, and the silhouette shifts. The prompt is not prose to be polished — it is a specification.
When you review a render, check identity before anything else: face shape, hairline, wardrobe, hands. If identity slips, regenerate with the same prompt before rewriting it, then tighten the reference input. Changing two variables at once is how creators lose track of what actually fixed the problem.
A useful habit is the three-pass review. First pass, identity and wardrobe only. Second pass, motion and physics. Third pass, sound and color. Trying to judge everything at once produces vague notes like that one looks off, which are useless when you open the project next week.
Camera, lens, and movement decisions per shot
Lighting and camera behavior are where AI footage either feels cinematic or feels like a screensaver. Decide both on paper, before prompts, so the decision is creative rather than accidental.
For lighting, specify direction, quality, and contrast. Soft key from the left window, warm, low contrast is a plan. Good lighting is a wish. Note the time of day and any practical sources — neon signs, lamps, phone screens — that will motivate the light inside the scene. Practical sources are especially valuable because they give the model something concrete to render.
For lens, think in emotional terms rather than technical ones. Wide lenses with deep focus make spaces feel large and characters feel small, which suits loneliness, scale, and comedy built on a character being outmatched. Long lenses with shallow focus isolate faces and compress backgrounds, which suits intimacy, suspicion, and quiet observation. Write the intended feeling in your notes, then translate it into prompt language such as wide angle, deep focus or telephoto compression, blurred background.
For camera movement, match the move to the dramatic function:
- Static — observation, stillness, dread, or comedy that depends on timing.
- Slow push in — growing realization or emotional pressure.
- Pull out — isolation, scale reveal, or closing a scene.
- Lateral tracking — travel, parallel action, or following a walking character.
- Orbit — hero shots, product beauty shots, reveals.
- Handheld — urgency, documentary realism, anxiety.
- Crane or drone rise — scale, finality, establishing geography.
One movement per shot. This rule sounds restrictive and is actually liberating. Two movements in a short clip usually produce mush, because the model splits its attention and completes neither move convincingly. If the story truly needs a compound move, split it into two shots and cut between them.
Finally, plan pacing deliberately. A sequence of six push-ins in a row feels monotonous even when each shot is beautiful. Alternate static and moving shots, alternate wide and close, and place your most dynamic move where the story earns it rather than where the render happens to look best.
Matching each shot to the right generation method
Not every shot deserves the same technique. Choose per shot using six criteria: subject complexity, motion complexity, duration, text and logo handling, style fidelity, and retry cost. Write the choice into your shot list so you are not improvising under time pressure.
Image-to-video is the workhorse for sequences with recurring characters or locations, because it begins from a frame you already control. When you lock the first frame with a location plate, the shot rarely has to invent its own world.
Text-to-video is strongest for landscapes, abstract transitions, establishing shots, and anything without a fixed identity. It is also the cheapest way to explore, so use it for tests before committing to a heavier approach.
Video-to-video and stylization passes are ideal for restyling reference footage or pushing a consistent look across a finished cut. If you already have live action, this path is faster than rebuilding the scene from scratch.
Match difficulty to method with a complexity budget. A shot containing two people interacting, a moving camera, and dialogue is three hard problems stacked. Either simplify the shot or accept many retries. The most reliable pattern is to reserve complicated content for simple framing and give ambitious camera work to shots with a single subject.
Two habits make this practical. First, never let a shot depend on a capability you have not already gotten working once; test risky shots early at low resolution and short duration. Second, build a small internal library of settings that work for your style — portrait close-up, walk-and-talk, vehicle motion, product turn, drone reveal — and reuse them rather than re-deriving settings every session.
Assembly and review: passes, triage, and versioning
Edit early and edit rough. Place your best takes on a timeline with temporary music, then review the sequence muted, then with sound. Muted viewing exposes problems the audio hides: unclear geography, repeated framing, pacing that sags in the middle, and cuts that match on nothing.
Classify failures instead of reacting to them individually. Typical failure types are identity drift, temporal flicker, distorted hands or props, unnatural physics, camera jitter, and jumpy cuts caused by mismatched eye lines. Each has a remedy ladder, and you should work down it in order:
- Choose a different take from the same batch.
- Re-render with the identical prompt and a fixed seed.
- Tighten the prompt by removing moving parts.
- Shorten the shot so the flaw falls outside the window.
- Switch generation method or model.
- Replace the shot with a simpler one that serves the same story function.
Most creators jump straight to step five, spend their remaining time there, and never discover that a simpler answer existed two steps earlier.
Version your files obsessively and name shots identically in the shot list, the board, and the timeline: sc02_sh014_v03_i2v.mp4. When a note arrives three weeks later asking for a change, you will know exactly which file to open. Archive raw renders instead of overwriting them; a rejected take becomes useful the moment the edit changes shape.
Structure folders before the project starts: script, shot list, boards, references, renders, selected, sound. Keep reference images for characters and locations in one place with descriptive filenames. When someone joins the project, hand over three things — the shot list with priorities and status, the board as a single contact sheet, and the continuity sheet. A contact sheet communicates an entire sequence at a glance, which makes client reviews dramatically shorter.
Mistakes that quietly ruin AI sequences
- Storyboarding finished-looking frames. Attachment to pretty boards leads to defending shots that do not cut.
- Rewriting descriptions per shot. Paraphrasing breaks identity faster than any model limitation.
- Skipping the function test. Shots without a purpose make a sequence feel long even when the runtime is short.
- Stacking two camera moves in one clip. The model completes neither move well.
- Mixing aspect ratios mid-project. Composition breaks the moment you crop.
- Generating before lighting is decided. Relighting in post is not a dependable repair.
- No priority column. When everything is critical, nothing is cuttable and the deadline decides for you.
- Testing ambitious shots last. Risky shots belong first, when a redesign is still affordable.
- Reviewing everything at once. Separate identity, motion, sound, and color passes.
- Chasing a spectacular shot that does not fit. A well-planned shot rendering at eighty percent of your vision usually serves the film better than a showstopper that breaks the sequence.
FAQ and a final pre-generation checklist
Do I need paid planning tools to work this way? No. A text editor, a spreadsheet, and any image generator can produce the four artifacts described here. Specialized tools mainly save time and automate consistency checks.
How long should pre-production take for a one-minute piece? Roughly two to four hours the first time, and well under an hour once the workflow is habitual. The savings appear later as fewer failed renders and faster edits.
Is a storyboard really mandatory? It is the cheapest place to make decisions. Skip it and you will make the same decisions during editing, when each one costs a render instead of a minute.
How many takes per shot should I generate? Three to five for simple shots, more for complex ones. If you need ten, the shot is usually misdesigned rather than unlucky.
What if a character will not stay consistent? Reduce complexity: one character per shot, static camera, identical wardrobe wording, and image-to-video anchored to a locked reference frame.
Can I mix models and tools within one project? Yes, and you usually should, because different shots have different demands. Just standardize color grading and grain at the end so the cuts feel like one film.
How do I know when a shot is done? When it serves its function and does not break continuity. Perfection is not the bar; coherence is.
Before you generate final footage, confirm: the shot list has functions and priorities, every character and location has a written description plus a reference image, each shot has one lighting plan and at most one camera move, the aspect ratio is fixed, and your riskiest shot type has already succeeded once at low cost. If all of that is true, generation stops being a gamble and becomes execution — which is exactly what a planning layer is for.


