Why Pre-Production Decides the Outcome of AI Video Work
Every week someone opens a text-to-video tool, types a paragraph, and waits. What comes back is often striking: beautiful light, interesting camera drift, a face that looks almost real. Then they try to build a ninety-second piece out of that footage and the illusion collapses. The clips share no color palette. The character's jacket changes material between shots. Nothing cuts together, because no two shots were ever designed to sit next to each other.
The failure is almost never the model's fault. It is a planning failure. Video generation is a rendering step, and like any rendering step it needs a plan to render. In traditional production that plan lives in two places: the script, which decides what happens and in what order, and the storyboard, which decides how each moment looks and where the camera sits. Those two artifacts are also the cheapest things you will produce all project and the most expensive things to skip.
AI has not removed the need for them. It has compressed the time they take. A structured script becomes a shot list in twenty minutes. A shot list becomes storyboard frames in an afternoon. Approved frames become moving clips in a few hours. Work that used to occupy a week of pre-production now fits into a single day, which is exactly why there is no longer a good reason to jump straight from a vague idea to a prompt.
This guide describes a neutral, tool-agnostic pipeline you can run with whichever generator, editor, and voice tool you prefer. It is organized around artifacts rather than features, because artifacts are what make a workflow repeatable and portable. If you switch tools next quarter, your script, shot list, and style bible come with you. Only the rendering step changes.
The two questions to answer before you generate anything
First: what is the single idea this piece must land? If you cannot write it in one sentence, the piece will drift. Second: who is watching, and where? A vertical clip watched on a phone with the sound off has completely different constraints than a landscape film watched on a laptop with headphones. These two answers shape runtime, framing, caption strategy, and pacing, and they cost nothing to settle in advance.
Why the debugging argument matters more than the quality argument
People usually frame pre-production as a quality measure. The stronger case is debuggability. When something looks wrong after generation, a structured pipeline tells you exactly where to look. A bad clip points back to a bad frame. A bad frame points back to a bad shot list entry. A bad shot list entry points back to an unclear script line. Without those layers, every problem forces you to start over from the idea, and starting over from the idea is how a weekend project turns into an abandoned folder.
The Seven Artifacts of a Reliable AI Video Pipeline
A dependable pipeline produces seven artifacts, each consumed by the next stage. Keep them in one project folder with predictable names.
- Concept brief โ one paragraph stating audience, tone, target runtime, delivery format, and the single idea the piece must land.
- Script โ scene-by-scene description with action, dialogue, and explicit visual intent.
- Shot list โ the script broken into individual shots with duration, framing, camera movement, subject action, and a visual note.
- Style bible โ reference stills, palette, lens and light language, and short physical descriptions of recurring subjects.
- Storyboard frames โ one approved still image per shot, generated in order and curated immediately.
- Motion pass โ approved frames animated with image-to-video, plus a short motion prompt per clip.
- Assembly โ edit, sound design, color treatment, captions, and delivery exports.
The design principle is that each stage has a single, cheap artifact. If a shot turns out wrong during animation, you go back to the frame, not to the idea. If a frame is wrong, you go back to one row of the shot list. Debugging becomes surgical instead of total.
Match pipeline weight to project length
A fifteen-second social spot needs a script and a rough shot list, and it can often skip a formal style bible if every frame is generated in one session with a fixed style clause pasted into each prompt. A three-minute brand film needs all seven artifacts. A ten-minute narrative short needs all seven plus a continuity tracker that logs costumes, props, time of day, and screen direction per scene. Time spent on structure scales with runtime, because continuity errors get more expensive the longer the piece runs.
A note on file naming
Adopt shot-014-frame-v2.png and shot-014-motion-v1.mp4 from the beginning. When you have eighty files, sortable names are the difference between a working folder and a scavenger hunt. Lock the numbering to the shot list and never renumber mid-project.
Writing a Script the Pipeline Can Execute
A script written for a human crew assumes that a cinematographer, an actor, and a location scout will fill in the gaps with judgment. An AI pipeline has no such collaborators, so anything left implicit may be invented badly. The fix is not to write more. It is to make visual intent explicit.
Keep scenes small and self-contained
Write in scenes that can be captured in one location with one lighting setup. A scene that jumps from a rooftop at dusk to a subway car at noon forces a hard reset of style references mid-sequence, and that reset usually shows up as a visible jump in look and character features. Where the story genuinely needs multiple locations, number them as separate scene blocks so the shot list inherits clean boundaries.
Separate action, dialogue, and visual notes
Use a light three-line discipline even if you are not writing in formal screenplay format:
- Action โ what physically happens on screen, present tense, one or two sentences.
- Dialogue or voice-over โ the words, with the speaker labeled.
- Visual note โ lens, light, color, and mood cues the storyboard stage must honor.
Visual notes are the highest-leverage lines in an AI pipeline. "Wide shot, cold blue morning light, wet asphalt reflections, slight handheld drift" gives the storyboard stage something concrete to translate. "Beautiful and cinematic" gives it nothing to work with.
Write to the target runtime
Estimate roughly two and a half to three seconds of finished screen time per line of action, and about a second and a half per short spoken sentence. A sixty-second piece therefore needs somewhere between twenty and thirty distinct beats. If your draft has sixty beats, you have written a three-minute film. Cutting beats on paper costs nothing; cutting shots after generation costs hours.
Name every recurring subject
Give each character, product, and location a stable name and reuse it verbatim. Do not call the same person "the woman," then "our protagonist," then "she." Stable names make it far easier to keep reference images and prompts aligned across dozens of shots, and they reduce the risk that a model treats two mentions of one person as two separate people. Use the name in the shot list too, so the frame stage inherits it automatically.
Write the last line first
A useful trick for short pieces: write the final beat, then write backwards. Endings are where AI video projects get vague, because a model will happily generate a moody slow-motion shot that means nothing. If you know the piece ends on a specific action or image, every earlier beat can be trimmed against it.
Turning the Script into a Shot List
A shot list is where the script becomes technically executable. Each row should contain shot number, scene, duration, framing, camera movement, subject action, and the key visual note. Build it in a spreadsheet so you can sort, filter, and reorder without breaking numbering.
Use coverage patterns instead of inventing every shot
Coverage is the practice of covering one scene from several angles so the edit has options. Three patterns do most of the work:
- Establish, engage, emphasize โ a wide to set the space, a medium to carry the action, a close-up for the emotional or informational beat.
- Two-shot plus singles โ for dialogue, a shared frame then alternating singles, with an insert for important objects.
- Move-match chain โ one continuous camera gesture split across two or three shots that the editor can join seamlessly.
Reusing patterns keeps the piece readable and speeds up prompting, because each pattern maps to a repeatable prompt template you can fill in rather than compose from scratch.
Do the duration math before generating
Short-form video lives or dies on pace. If your shot list totals ninety seconds of generated footage for a forty-five-second edit, you have a healthy ratio. If it totals exactly forty-five seconds, you have no room to trim around awkward motion, which is the single most common reason AI clips feel sluggish. Aim for roughly one and a half to two times coverage on action-heavy pieces and about one and a third on dialogue-driven pieces.
Flag the shots generation handles badly
Some shot types are consistently harder: complex hand interactions, crowds, legible text on screen, rapid direction changes, and characters eating or drinking. Where the story allows, replace these with framing that reads the same idea more simply. A close-up of a hand lifting a mug carries a breakfast scene better than a full-body wide. Flagging risky shots in the shot list lets you budget extra retries for them instead of discovering them at the end of the schedule.
Assign a screen direction to each scene
Decide early whether a character moves left-to-right or right-to-left across the frame within a scene, and keep it consistent. Audiences read direction as geography, and flipping it mid-scene makes two shots feel like they are in different places. This is a one-column addition to the shot list that prevents an entire category of confusing edits.
Building a Style Bible Before Generating Frames
The style bible is the smallest artifact in the pipeline with the largest effect on consistency. It contains reference stills, a palette, a lens and light language, and one-line physical descriptions of recurring subjects.
Reference frames do the heavy lifting
Collect four to eight stills that capture the look you want: color, contrast, grain, depth of field. You are not copying them; you are describing a target. Turn them into a written style block of twenty-five to forty words that you paste into every image prompt. Something like: "muted teal and amber palette, soft overcast key light, 35mm lens character, shallow depth of field, subtle film grain." Consistency improves more from a fixed style block than from a fixed random seed alone, because the block survives across tools and sessions.
Character sheets prevent drift
For each recurring subject, generate a small sheet: a neutral front-facing portrait, a three-quarter view, and a full-body shot in the primary costume. Keep the descriptions short and physical โ age range, hair, build, one or two distinguishing features, wardrobe. Avoid subjective adjectives like "magnetic" or "world-weary." Models cannot render mood words into consistent features, but they respond well to "short cropped black hair, scar above the left eyebrow, olive field jacket."
Lock aspect ratio and resolution first
Decide delivery format before generating a single frame: 16:9 for landscape, 9:16 for vertical social, 1:1 or 4:5 for feed placements. Recomposing a finished storyboard from landscape to vertical is not a crop; it usually requires regenerating frames with different framing. Choosing early saves an entire second pass, and it also determines how much headroom you leave on subjects.
Keep a continuity log
Alongside the style bible, keep a running list of facts: whose jacket is unbuttoned, which side of the car the driver sits on, what time of day it is in scene four. Update it as you generate. A ten-line log prevents the most embarrassing class of errors, where a prop or a costume changes between two shots that are supposed to be seconds apart.
Generating and Curating Storyboard Frames
Now the pipeline becomes visual. Work shot by shot, in order, and treat generation as casting plus blocking plus framing in one step.
Build a reusable frame prompt template
A dependable template has five parts in a fixed order: subject and action, wardrobe and props, environment and time of day, camera framing and lens, and the style block. Keeping the order constant makes it much easier to diagnose a bad result, because you can change one clause at a time instead of rewriting everything. Save the template as a text snippet and duplicate it for each row of the shot list.
Curate hard, and curate immediately
Generate three to five candidates per shot and pick one right away. Do not save maybe frames. Storyboard stages balloon when a folder fills with near-duplicates, and the real cost is not storage โ it is decision fatigue later, when you cannot remember why a frame was kept. Delete the rejects and rename the approved frame with its shot number so the motion stage can process it in a predictable order.
Read the sequence as a silent film
Place approved frames in order and scroll through them quickly with no sound. If the story is not legible at that speed, the problem is in the shot list, not the frames. Fixing framing at the storyboard stage costs minutes. Fixing it after motion generation costs hours, plus the money and time spent on the clips you throw away.
Judge frames by what they set up
A frame is not a poster. Its job is to make the next frame make sense and to give the editor a clean place to cut. A technically beautiful frame that duplicates the information of the shot before it is dead weight. Ask of every frame: what new information does this add, and what does it set up?
Animating Approved Frames into Clips
Image-to-video generation is the most controllable approach for narrative work because composition is already decided. Your job at this stage is to add motion that respects the frame instead of fighting it.
Keep motion prompts short and physical
Describe one primary motion and, at most, one secondary motion. "Slow push-in as she turns her head toward the window" is workable. A prompt listing four simultaneous movements usually produces warping, because the model has to reconcile motion vectors that conflict in the same pixels. If a shot needs complex action, split it into two shots and let the edit create the sense of continuity.
Choose camera moves that match the emotional beat
- Push in โ increasing focus, intimacy, or tension.
- Pull out โ revealing context, isolation, or scale.
- Lateral track โ establishing space and the relationship between subjects.
- Orbit or arc โ energy, product hero moments, reveals.
- Static with subject motion โ dialogue and detail shots, where camera movement adds nothing.
Mixing too many move types in a thirty-second piece makes it feel like a demo reel rather than a film. Pick two dominant moves and use them consistently. A piece can absolutely be built on static frames plus subject motion alone, and it will often look more confident than one that drifts constantly.
Generate slightly longer than you need
Ask for four to six seconds when the edit needs two to three. Extra frames give you handles for cutting on motion and for trimming the first and last frames, where generated motion most often looks unnatural. Trim the entry and exit and most clips will cut together cleanly, even when the middle of the clip has a small imperfection.
Check every clip against the artifact list
Before approving a clip, look for face morphing during head turns, hands that change shape mid-gesture, background elements that swim or duplicate, clothing that shifts texture, and shadows that move independently of their subject. Most of these reduce if you shorten the clip, simplify the motion prompt, or reduce how much of the subject moves at once. Sometimes the fastest fix is to regenerate the source frame with a simpler pose.
Holding Consistency, Sound, and Delivery Together
Consistency is the difference between a collection of clips and a film. It operates on four levels: character, environment, color, and motion language.
Consistency across four levels
Character. Reuse the same reference image and the same descriptive clause for every appearance. Where the tool supports subject referencing, use it instead of relying on text alone. If a costume changes during the story, create a separate reference for each costume state and treat them as distinct subjects.
Environment. Keep a location reference frame for every recurring set. Describe props that stay put โ the red kettle, the chipped blue door โ and mention at least one in every shot set in that location. Repeated anchors make continuity feel intentional rather than accidental.
Color. Apply one grade across the whole piece rather than grading each clip independently. A single consistent treatment makes small differences in generated color temperature stop being noticeable.
Motion. Decide the overall camera energy: locked-off, gentle drift, or handheld. A handheld documentary piece and a locked-off product piece can both look excellent, but a piece that alternates randomly between them looks unfinished.
Build audio in layers
Generated picture without sound design reads as a technical demo. Sound is where AI video starts to feel like film. Work in four layers: voice (record real narration where possible, and never mix two synthetic voices for the same character), ambience (one continuous room tone or environment bed per location, crossfaded at scene changes), foley (footsteps, cloth, doors, object handling โ these carry more perceived realism than music), and music (choose a tempo that matches your cut rhythm and cut to the beat where it helps).
Edit for rhythm, not completeness
Cut on motion, not after it. When a subject begins to move, cut a few frames into the movement. This hides generated motion imperfections and makes the piece feel deliberate. Most AI footage benefits from being cut fifteen to twenty percent tighter than the first assembly suggests. If a clip drags, try trimming its head before its tail; entrances usually contain more dead time than exits.
Delivery checklist
Confirm aspect ratios per placement, loudness normalization targets for each platform, burned-in captions versus caption files, and safe areas for vertical formats where interface elements cover the top and bottom of the frame. Export a master with no captions plus platform-specific versions, so you never have to re-edit to fix a caption typo.
Common Mistakes and Tool Selection Criteria
The same handful of problems show up in almost every struggling AI video project. Recognizing them early saves entire evenings.
- Skipping the shot list. Symptom: clips that look good individually but cannot be edited together. Fix: write the shot list before generating anything, even a rough one.
- Over-writing prompts. Symptom: unpredictable results that are hard to reproduce. Fix: cap prompts at one subject action, one camera move, and a fixed style block.
- Chasing a perfect single clip. Symptom: hours lost on one shot. Fix: set a retry limit of three attempts, then change the framing or simplify the action.
- Ignoring duration math. Symptom: an edit that drags. Fix: generate one and a half to two times coverage and cut to the tightest version.
- Inconsistent characters. Symptom: the lead looks like a different person in every scene. Fix: lock one reference image and one description clause per character and reuse them word for word.
- No sound pass. Symptom: the piece feels like a slideshow. Fix: add ambience and foley before adding music.
- Generating in the wrong aspect ratio. Symptom: a beautiful landscape storyboard you cannot use vertically. Fix: decide delivery format first.
Decision criteria for each tool slot
You do not need one platform that does everything. Most strong pipelines combine three or four specialized tools, and the criteria are straightforward.
- Script and outlining tools โ choose on structural support and clean export, not prose quality. You will rewrite the prose anyway.
- Image generation โ choose on subject referencing and control over framing and aspect ratio. Consistent character features matter more than raw photorealism.
- Image-to-video โ choose on motion coherence and clip length. Test the same frame across candidates and compare how faces and hands hold up.
- Voice โ choose on delivery consistency and commercial licensing terms.
- Editing โ any capable nonlinear editor works. What matters is that you can grade, mix, and caption in one place.
Where to invest your time
Invest in whichever stage currently produces your most frequent rework. If you constantly regenerate frames, your style bible is too thin. If you constantly regenerate motion, your frames contain too much implied movement or your motion prompts are too ambitious. If you constantly re-cut, your shot list lacked coverage options. If you constantly fix character continuity, you are paraphrasing your character description instead of pasting it.
Frequently Asked Questions
Do I need a formal screenplay format?
No, but you need consistency. Scene blocks with an action line, a dialogue line, and a visual note line cover everything the later stages require. Format matters less than the discipline of writing visual intent down somewhere the frame stage will read it.
How many storyboard frames should I generate per shot?
Three to five candidates for complex shots, one to three for simple inserts. Pick immediately and delete the rest. If you find yourself generating ten candidates for the same shot repeatedly, the shot is probably too complicated for the piece and should be simplified or split.
Can I generate video directly from a script and skip storyboards?
For abstract or purely atmospheric pieces, yes. For anything with characters, dialogue, product detail, or a narrative arc, storyboards pay for themselves within the first few retries. The storyboard is also where you discover that a shot is impossible before you have spent time animating it.
How long should individual clips be?
Generate four to six seconds and edit down to two or three. Handles make clean cuts possible and cost little extra time. Very short generations sometimes produce almost no motion at all, while very long generations tend to accumulate drift, so the middle range is the sweet spot.
What is the fastest way to fix an inconsistent character?
Reduce the number of shots featuring them, lock one reference image, and reuse an identical description clause in every prompt. Adding more detail rarely helps; repeating the same detail does. If the character still drifts, check whether your style block is fighting the character description.
Is vertical video just a crop of the landscape version?
Usually not. Plan vertical framing separately, favoring closer shots and centered subjects with headroom for interface elements. Some shots translate well with a crop, but wide establishing shots almost never do, because the information lives at the edges of the frame.
How do I know when a shot is good enough?
Ask whether it works silently, at speed, in sequence. If it does, it is good enough, regardless of how it looks in isolation. Perfectionism on individual clips is the most common way a short project becomes a long one.
What if my tools change halfway through a project?
Because the pipeline is organized around artifacts rather than features, this is survivable. Your script, shot list, style bible, and approved frames all transfer. Only regenerate the motion pass with the new tool, and compare a single test clip before committing to a full re-render.
Putting the Pipeline to Work
A dependable AI video pipeline is less about the model you choose and more about the artifacts you build between idea and render. A script that states visual intent, a shot list with durations and coverage, a style bible that fixes look and character, curated storyboard frames, motion prompts that respect the frame, and a sound pass that gives the picture weight โ these are the levers that separate a finished piece from a folder of attractive clips.
Start with one short project and build all seven artifacts, even in rough form. Keep them in a single folder with predictable names. Read the frames as a silent sequence before you animate anything. Cut on motion. Paste your character description instead of paraphrasing it. After the second or third project, the templates stabilize and the pipeline stops being a checklist and becomes a rhythm: write, plan, frame, animate, cut, mix, deliver. That rhythm is what makes AI video production repeatable instead of lucky.




