A screenplay that reads beautifully on the page can fall apart the moment you try to generate it. The problem is rarely the story. It is the gap between prose written for humans and instructions a video generator can actually execute. Closing that gap is the real craft of AI filmmaking, and it starts long before you type your first prompt.
Why Script-First AI Video Beats Prompt Roulette
Most people who try to make a film with generative video begin with a beautiful prompt and end with a folder of unrelated clips. Each shot looks impressive on its own. Put three of them end to end and the illusion collapses: the lighting jumps, the character's jacket changes color, the camera language has no logic, and nothing cuts together because nothing was designed to cut.
Prompt-first workflows fail for three predictable reasons.
Tonal drift. Without a written reference point, each prompt session drifts a little. Shot two is warm and intimate, shot nine is cold and epic, and the film has no single voice. A script pins down tone before any pixels exist.
No coverage. Editors need options. A script that has been broken into a shot list tells you that a tense conversation needs a wide, two singles, and a close-up on hands — so you generate all four and choose in the edit. Prompt-first creators generate one version of the scene and then discover they have no way to build tension.
Continuity breaks. Costume, time of day, weather, props, and geography must be consistent across every shot. That is a production-design problem, not a prompt problem. It gets solved on paper first and enforced second.
Traditional production solved these problems decades ago with a simple sequence: script, breakdown, shot list, storyboard, then shoot. AI video has not changed that logic. It has only changed who can afford to follow it — now that is basically anyone with a laptop and a clear plan.
The Five-Stage Pipeline From Page to Picture
Every reliable AI film follows the same five stages, whether it is a nine-second social clip or a ten-minute brand documentary.
- Script for generation — write a screenplay whose scene headings, action lines, and dialogue are already half-way to a shot description.
- Breakdown and shot list — convert scenes into individual shots with purpose, duration, and camera intent.
- Visual bible — lock characters, palette, lens language, and locations as reusable references.
- Per-shot generation — prompt, preview, and refine each shot with motion and continuity controls.
- Assembly and finishing — edit for rhythm, add sound design and music, grade, caption, and deliver.
Two things make this pipeline work in practice. First, the loop: you will return to the shot list constantly as shots prove harder or easier than expected. Second, the order: never generate a final-quality shot before the script and shot list are frozen. A reshoot in AI is free in theory but expensive in time, because you must rebuild continuity every time the plan moves.
Stage 1: Writing a Screenplay AI Can Actually Shoot
You do not need to write differently in terms of story. You do need to write differently in terms of description density.
Scene headings carry real information
"INT. BAKERY — EARLY MORNING — RAIN" gives a generator four facts: interior, location type, time of day, weather. That single line becomes the backbone of ten prompts later. Put time of day and weather in every heading, even when it feels redundant, because those are the two variables AI generators most often invent for themselves.
Action lines: nouns over adverbs
Vague action is the single biggest cause of wasted generation time. Compare:
- "She walks in, looking nervous." — nervous is an internal state; the generator cannot render it.
- "She pushes the door with two fingers, shoulders raised, eyes darting to the counter." — three visible, renderable behaviors.
Rewrite every action line until a stranger could draw it. Physical detail is also your continuity anchor: once you write "shoulders raised," you know what to keep constant across shots.
Dialogue is a performance brief
Unless you are making a subtitled silent film, spoken dialogue needs a delivery: pace, volume, physical action, and eyeline. Mark those in the script. If your pipeline cannot yet produce convincing lip sync for a given shot, write the scene so the line is delivered off-camera, over the shoulder, or as voice-over during a cutaway. That is not cheating — it is screenwriting that respects the medium.
Design for your shot vocabulary
If you know your tool handles slow dolly-ins beautifully and crowds badly, write fewer crowds. A script that demands impossible shots is a script that will be rewritten in the edit anyway. Skim your draft and tag every scene with a difficulty note: easy, medium, risky. Then budget your time toward the risky ones.
Stage 2: Turning the Script into a Shot List
A shot list is where a screenplay becomes a production plan. Keep one row per shot and add the fields that matter for generation.
| Field | Why it matters |
|---|---|
| Shot ID | Naming convention for files and versions (s03_sh02_v4) |
| Purpose | Establishing, reaction, insert, transition, payoff |
| Duration | Target seconds; drives how much motion you can afford |
| Framing | Wide, medium, close, macro, over-the-shoulder |
| Camera move | Static, push in, pull out, pan, handheld, orbit |
| Continuity tokens | Costume, prop, time of day, weather, location |
| Audio | Dialogue, ambience, music cue, silence |
A practical rule: aim for roughly three to five times more shots than scenes. A 60-second film with eight scenes usually lands at 18–24 shots, averaging 2.5–3.5 seconds each. Short shots hide small inconsistencies; long shots expose them.
Also mark which shots are hero shots — the four or five frames that carry the story. Spend your refinement time there and accept "good enough" on the connective tissue.
Stage 3: Building a Visual Bible for Consistency
Consistency is the hardest problem in AI video, and it is solved with references, not adjectives. A visual bible is a small document with four parts.
Character sheets. For each principal, create a reference image (or a small set: front, profile, three-quarter). Write a fixed description string you will paste into every prompt without editing: age range, hair, build, wardrobe, distinguishing details. Never improvise this wording mid-project — small wording changes produce different faces.
Palette and light. Pick two or three colors and one dominant lighting condition per location. "Teal shadows, warm practicals, overcast daylight" is a rule you can enforce across thirty shots.
Lens language. Decide what your film looks like optically: shallow depth of field for intimacy, wide-angle for unease, long-lens compression for crowds. Assign a lens feel per scene type and keep it stable.
Location anchors. Two or three still frames per location, generated early and reused as image references. This is what stops a kitchen from having a different window in every scene.
The payoff is mechanical: instead of describing your film from scratch in every prompt, you assemble prompts from locked components.
Stage 4: Prompting Shot by Shot
A dependable prompt has eight slots. Fill them in the same order every time.
- Subject — who or what, using the exact character string from the bible
- Action — one visible verb phrase, present tense
- Environment — location, time of day, weather
- Camera — framing and movement
- Lens — focal feel, depth of field
- Light — source, direction, quality
- Motion — what moves and how fast
- Style — film stock, grade, genre reference
A filled example: "Mid-30s woman in a mustard raincoat, pushing open a glass bakery door with two fingers, small bakery interior, early morning, heavy rain outside, medium shot slowly pushing in, 50mm shallow depth of field, soft overcast window light from the left, steam drifting from a coffee cup, muted teal-and-amber grade, documentary realism."
Two habits separate fast workflows from slow ones.
Preview at low resolution first. Generate a draft pass for every shot in the film before refining any single shot at full quality. You will discover pacing and continuity problems that are invisible in isolation.
Reuse seeds and references. When a shot is close but not right, change one variable — motion, lens, or light — and regenerate with the same seed and reference image. Changing three variables at once teaches you nothing about what worked.
Control the first and last frame. For shots that must match a neighbour, generate the boundary frames first and use them as start and end conditions. This is the closest thing AI video has to matching a cut.
Stage 5: Assembly, Sound, and Finishing
AI video gets the attention; the edit gets the result. Import every approved shot into a timeline and cut for rhythm before you cut for logic. Silent sequences usually need to be 10–20 percent shorter than the script suggested, because generated motion reads slower than performed motion.
Then build sound in layers:
- Ambience first, so the world feels continuous across cuts
- Hard effects on visible actions — doors, footsteps, fabric, glass
- Dialogue and voice-over, cleaned and levelled
- Music last, chosen to match the emotional arc rather than the genre
Finish with a single grade pass to unify color across shots, a subtle grain or film texture to hide minor generation artifacts, and burned-in or platform captions. Export one master and separate aspect-ratio versions: 16:9 for web and presentations, 9:16 for vertical feeds, 1:1 for thumbnails.
Matching the Generator to the Shot
Different shot types reward different tool strengths. Rather than committing to one engine, keep a short list and route each shot.
| Shot type | What to look for |
|---|---|
| Talking head | Stable identity, natural head motion, audio-driven sync |
| Product macro | Sharp texture, controlled reflections, slow push |
| Establishing wide | Coherent depth, believable sky and horizon |
| Action | Motion blur handling, no limb warping |
| Stylized/illustrated | Strong style adherence, consistent line weight |
| Insert/detail | Fine control of first and last frame |
Test each candidate tool on the same three prompts before a project starts. Keep a note of which tool handled which shot type best; that note becomes your routing table for every future film.
Mistakes That Quietly Eat Your Week
- Overlong prompts. Beyond roughly 60–80 words, extra detail competes with itself. Cut adjectives before you cut structure.
- Changing style words mid-project. "Cinematic" in shot three and "moody film still" in shot twelve produces two different films.
- Generating finals too early. Full-quality passes on an unfinished plan guarantee rework.
- Ignoring aspect ratio until export. Vertical and horizontal framing are different compositions, not crops.
- No naming convention. Without
scene_shot_versionfilenames, you will edit the wrong take. - Assuming sync is solved. Plan dialogue shots with fallback coverage so a failed sync does not break the scene.
A Worked Example: 45-Second Brand Film
A coffee roastery wants a 45-second product film. The brief: craft, warmth, morning ritual.
Script: six scenes, 14 lines. Shot list: 20 shots — four establishing, six product inserts, five human moments, five connective transitions. Visual bible: two characters, one location, palette of warm amber and deep walnut, 35mm and 50mm feel, overcast window light with practical warmth.
Previz: every shot generated at draft quality, cut together the same afternoon. Two problems surfaced immediately — the opening establishing shot was too slow, and a mid-film insert broke the palette. Both were fixed in the shot list, not in the edit.
Finals: hero shots (the pour, the first sip, the closing window frame) regenerated at full quality in three rounds each. The remaining 17 shots needed one or two passes.
Finish: ambience of rain and grinder hum, music that lifts at the pour, 20 percent warm grade, three exports. Total elapsed time: two focused days.
Pre-Publish Quality Checklist
- Story is understandable with the sound off
- Every cut has a reason — no shot is there because it looked nice
- Character wardrobe, hair, and props match across scenes
- Time of day and weather are consistent within each scene
- Camera language does not change without intent
- Audio levels are even; no clip is louder than the rest
- Captions are accurate and legible on mobile
- Correct aspect ratio and safe margins for each destination
FAQ
Do I need a traditional screenplay format?
Not strictly, but scene headings and consistent action lines give you two things nothing else does: a shared reference for every collaborator and a clean basis for a shot list. Even a 400-word treatment in scene format works.
How long should each generated clip be?
Aim for 2–4 seconds for most shots. Longer clips are harder to keep consistent and rarely survive the edit. Reserve longer durations for a deliberate slow push or a held image that carries emotion.
What is the best way to keep a character consistent?
Lock three things and never change them: a reference image, an exact written description string, and a wardrobe rule. Consistency is a discipline problem more than a model problem. If the face keeps shifting, your prompt wording is probably drifting.
Should I write dialogue if my tool cannot do lip sync well?
Yes, but stage it cleverly. Write lines for off-camera delivery, over-the-shoulder coverage, or voice-over during cutaways. Then add one or two close-up dialogue shots and test whether sync holds. If it fails, you have already shot the scene without it.
How many generations does one final-quality shot take?
Realistically three to eight, more for complex action. Plan for it by making the shot list shorter rather than accepting whatever comes out first. A 20-shot film finished well beats a 40-shot film finished badly.
Where does music fit in the workflow?
Choose it after the first rough cut, not before. Music written against a script rarely fits the pacing of generated shots. Cut the picture to rhythm first, then find a track that matches the arc you actually built.
Can this workflow handle a longer film?
Yes, with one change: work in scene blocks. Finish and lock one scene completely before moving to the next, and keep a running continuity sheet of wardrobe, props, and lighting states. Long-form AI video fails when continuity is managed from memory instead of from a document.


