Why AI Belongs in Preproduction, Not Just Post
For years, generative tools in video were treated as a finishing trick: upscale this frame, remove that boom mic, generate a background plate. The more interesting shift is happening earlier in the pipeline. Preproduction is the phase where an idea becomes a script, a script becomes a shot list, and a shot list becomes something a crew — or a render queue — can actually execute. It is also where most projects are quietly won or lost.
Language models happen to be good at exactly the work preproduction demands: restructuring text, holding constraints, producing variations on demand, and translating between formats. Image models can turn a written shot into a visual reference in seconds. That combination does not replace a director. It removes the friction between having an idea and being able to see it, which is the single biggest cause of stalled video projects.
This guide walks through a complete, tool-agnostic workflow: brief, treatment, script, shot list, storyboard frames, animatic, then handoff to a generative video model. It also covers the parts that usually break — character consistency, camera language, and the habit of accepting the first output a model gives you.
The Six Stages of an AI-Assisted Preproduction Pipeline
The temptation with AI tools is to skip straight to generation. Resist it. A structured pipeline produces far better raw material, and the structure itself is what makes revision cheap later.
Stage 1: Brief and logline
Start with a single paragraph that states the deliverable, the audience, the runtime, the format (vertical, widescreen, square), and the emotional target. Feed that paragraph to a language model and ask for ten logline variations at different angles: comedic, tense, documentary, absurd. You are not looking for the perfect line. You are looking for the angle that makes you want to keep working.
Stage 2: Treatment
Expand the chosen logline into a one-page treatment with a beginning, a turn, and a payoff. Ask the model to write it in present tense and to name the visual motif that repeats. That motif becomes your continuity anchor when you are generating frames three days later and cannot remember what the piece is about.
Stage 3: Script
Only now do you convert the treatment into a shooting script. Specify the format you want — scene headings, action lines, dialogue, parentheticals — and insist on a maximum runtime. A common failure is a treatment that reads well but expands into four minutes of script for a ninety-second video.
Stage 4: Shot list
Break the script scene by scene. For each scene, produce a table with shot number, description, shot size, camera movement, lens suggestion, duration, and audio note. This is the document that most people skip and most productions regret skipping.
Stage 5: Storyboard frames
Generate one still per shot using an image model, then assemble them into a numbered contact sheet. The goal is not beauty. The goal is to prove that the sequence reads, that the geography makes sense, and that a viewer can follow the action without dialogue.
Stage 6: Animatic
Drop the frames into an editor — DaVinci Resolve, Premiere Pro, or anything that gives you a timeline — hold each frame for its planned duration, and lay a scratch voiceover or temp music underneath. Watch it at full speed. Almost every pacing problem becomes obvious here, and fixing it costs nothing.
Writing Prompts That Produce Usable Scripts
Most disappointing AI scripts come from vague prompts. A model asked to "write a video about coffee" will produce generic copy because the request contains no constraints. Constraints are the entire craft.
Useful constraints to include:
- Runtime and word budget. Roughly 140–150 spoken words per minute for narration. A 60-second script is about 140 words, not 400.
- Point of view. Second person for instructional content, third person for narrative, first person for personal essays.
- Structure. Hook, context, escalation, payoff. Or problem, agitation, solution. Name it explicitly.
- Banned vocabulary. Tell the model which filler words to avoid — "unlock," "elevate," "in today's fast-paced world." This single instruction improves output more than any other.
- Reading level. Ask for short sentences and concrete nouns if the script will be read aloud.
- Format. Scripts for voiceover read better as one idea per line than as dense paragraphs.
A workable prompt pattern looks like this: role, deliverable, constraints, structure, and a single example of the tone you want. Keep the example short. Long examples cause the model to paraphrase rather than write.
Iterating without losing the thread
Ask for three variants of a single scene rather than three full scripts. Compare them, steal the best lines, and then ask for a rewrite that merges them. Working at scene level keeps the model anchored and keeps you in control of the overall shape.
Keeping Characters and Locations Consistent
Consistency is the hardest problem in AI-assisted video, and it starts in preproduction. If your storyboard frames show a character with different hair, different age, or different clothing in every shot, the final renders will drift even further.
Build a character sheet before you generate a single frame. For each recurring character, write a fixed description block and reuse it verbatim in every prompt: approximate age, build, hair, wardrobe, distinguishing feature, and one sentence of attitude. Do the same for locations — a fixed description block for the kitchen, the rooftop, the forest path.
Then convert those blocks into prompts with a consistent prefix. The order matters: subject first, then wardrobe, then environment, then lighting, then lens, then style. Models weight early tokens more heavily, so the subject should never be buried at the end of a long sentence.
If your image tool supports reference images, seed images, or character training, use them. A single approved reference frame reused across twenty shots will do more for consistency than any amount of prompt engineering. Approve one frame per character first, then generate everything else from it.
Handling wardrobe and time changes
If a character changes clothes or ages between scenes, treat it as a separate character entry with a shared name — for example, "Mara, Act One" and "Mara, Act Three." This prevents the model from averaging the two looks into something unrecognizable.
Shot Language Every Storyboard Should Carry
A storyboard without camera information is just an illustrated synopsis. Each frame should carry at least four pieces of technical intent.
Shot size. Wide, full, medium, medium close, close, extreme close, insert. Shot size controls emotional distance more than any other variable. Cutting from a wide to an extreme close-up is a statement; staying in medium shots is a neutral report.
Camera movement. Static, pan, tilt, dolly in, dolly out, truck, crane, handheld, gimbal, drone. In generative video, movement prompts are less reliable than in live action, so plan a fallback: any shot that a model renders poorly as movement can often be generated as a static frame and moved in the edit with a subtle push.
Lens and depth. Wide lens for spatial context, long lens for compression and intimacy, shallow depth of field for isolation. Mentioning a specific focal length in a prompt changes composition noticeably.
Lighting direction. Key from the left, backlit, practical lamps in frame, overcast diffusion, hard noon sun. Lighting is the fastest way to make a set of frames feel like one film instead of twenty unrelated images.
Add a duration estimate to each shot. Thirty seconds of screen time with only four shots will feel static; ten shots in fifteen seconds will feel frantic. Writing durations forces you to confront pacing before you render anything.
Translating Storyboard Frames Into Video Prompts
Once a frame is approved, it becomes the source of truth for the video prompt. A reliable structure:
- Subject and action — who is in frame and what changes over the duration.
- Environment — where they are, what is visible in the background.
- Camera — shot size and movement, described as a single clear instruction.
- Lighting and time of day.
- Style and texture — filmic, documentary, animated, archival.
- Duration and pacing — a slow reveal versus a quick beat.
Keep each prompt to one primary action. Models struggle when a prompt asks a character to walk, turn, and speak simultaneously while the camera also cranes upward. Split complex beats into two shots and cut between them.
If your tool supports image-to-video, always start from the approved storyboard frame. Starting from text reintroduces every consistency problem you already solved. If it supports motion or camera controls separate from the prompt, use those controls instead of describing motion in text — explicit controls beat adjectives almost every time.
Choosing Tools for Each Stage
You do not need one platform to do everything. A shortlist by function:
- Script and treatment: large language models with long context windows, plus a plain text editor for your own edits. Google Docs or Notion work fine for tracking versions.
- Shot lists: a spreadsheet or a structured Markdown table. Avoid anything that hides the table behind a UI when you need to paste it into a prompt.
- Storyboard frames: an image generator with reference-image support and a consistent seed workflow. Midjourney, Stable Diffusion variants, and Krea all handle this well.
- Animatic: any NLE. DaVinci Resolve is free and handles frame-hold timelines perfectly.
- Voice scratch: a text-to-speech tool or your own phone recording. Perfection is irrelevant; timing is everything.
- Video generation: pick based on the shot type you need most — some tools excel at realistic humans, others at stylized motion or landscapes.
Decide based on iteration speed, not on the best demo you have seen. A tool that produces a great frame in five minutes but takes twenty attempts is worse than a modest tool that lands in two.
Seven Mistakes That Sink AI Video Projects
1. Skipping the shot list. Without it, you generate attractive clips that do not cut together. You end up with a mood reel instead of a video.
2. Writing scripts for the page instead of the ear. Read every line aloud. If you run out of breath, the line is too long.
3. Letting the model choose the structure. Models default to generic three-act shapes. Impose your own.
4. Approving the first frame. The first generation is a draft. Generate at least four variants before locking a look.
5. Changing style mid-project. Every style change in a prompt needs to be applied retroactively to all previous frames, or the piece will look assembled from different films.
6. Ignoring audio in the shot list. Silence in the storyboard means a panic in the edit. Note where music drops, where narration pauses, where a sound effect carries the cut.
7. Rendering before the animatic works. If the story does not hold with still frames held for the right durations, it will not hold with expensive generated motion either.
A Pre-Render Quality Checklist
Before committing to a long render session, run through this list:
- Every shot has a purpose you can state in one sentence.
- No two consecutive shots have the same size and angle unless that repetition is intentional.
- Character description blocks are identical across every prompt in the project.
- The shot list includes durations, and the total matches your target runtime.
- The animatic has been watched at full speed at least twice.
- Narration has been read aloud and timed with a stopwatch.
- Audio cues exist for every transition.
- You have a fallback plan for any shot that depends on a model behaviour you have not tested yet.
- All prompts and approved frames are stored in one folder with numbered filenames.
That last item sounds trivial and saves entire days. Numbered assets let you rebuild a sequence after a tool update changes how it interprets prompts.
Building a Repeatable Personal Workflow
Once you have run this pipeline two or three times, template it. Keep a document with your prompt skeleton, your character description blocks, your style suffix, and your shot list columns. New projects then start from a filled-in template rather than a blank page.
Track which prompts worked and which produced unusable output. A simple log with the prompt, the tool, and a one-word verdict builds into a personal knowledge base that is more valuable than any general tutorial. Over a few months, you will notice patterns — certain phrasings reliably produce better lighting, certain shot types consistently fail in certain tools — and you can plan around them instead of rediscovering them.
Finally, keep a human in the loop at every stage. The value of these tools is not that they make decisions. It is that they make your decisions visible quickly, so you can throw them away and decide again.
Frequently Asked Questions
How long should a storyboard take with AI assistance?
For a one-minute video, expect two to four hours for a full pass: brief, treatment, script, shot list, frames, and animatic. Most of that time goes into reviewing and rewriting, not generating. A rough pass with placeholder frames can take under an hour and is often enough to validate the concept.
Can AI write a script that is actually good?
It can produce a structurally sound, competent draft quickly. What it cannot do is know what you find interesting. Treat its output as a first draft that reveals the shape of the piece, then rewrite the parts that carry meaning. The rewrite is where the quality comes from.
How many storyboard frames do I need?
One per shot in your shot list, plus a few alternates for key moments. A sixty-second video typically has twelve to twenty-five shots. If you cannot describe a shot clearly, it usually means the beat is not yet conceived properly, not that you need more frames.
Do I need a shot list if the AI generates video directly from a script?
Yes, unless your video has no cuts. The moment you have two shots, someone has to decide how they relate in size, angle, and duration. A shot list is the cheapest place to make those decisions.
What is the best way to keep a character's face consistent?
Lock one approved reference frame per character, reuse its description block word for word, and use image-to-image or reference-image features wherever they exist. Prompt wording alone is a weak lever compared to an actual reference image.
Should I generate video for every storyboard frame?
No. Generate only the shots that carry the story. Background shots, inserts, and transitions can often be produced as a still with a gentle push in the edit, which is faster and more controllable.
How do I handle a tool update that changes my results?
Keep your prompts and approved frames numbered and archived. When behaviour changes, re-test one shot per character against a known-good frame. If the gap is large, adjust the style suffix and note the change in your prompt log.
Where does AI fit least well in this pipeline?
Pacing and emotional judgement. Models can produce shots faster than you can judge them, so the bottleneck moves to your own taste. That is not a problem to solve — it is the part of the work worth protecting.



