Why AI Reshapes Pre-Production More Than Post
Most people assume generative video changes the editing suite. In practice, it changes the desk before the shoot. When a shot can be described in a paragraph and rendered in ninety seconds, the bottleneck moves upstream: what do you actually want to see, and can you describe it precisely enough for a model to reproduce it three times in a row?
That shift is why pre-production is now the highest-leverage phase of a short film. A director who spends two days writing a rigorous scene breakdown will out-produce a director who spends two days clicking through generations, even if the second director has better tools. The reason is simple: generative video rewards specificity, and specificity is a writing problem, not a rendering problem.
This guide walks through a complete pre-production workflow for AI-assisted short films. It covers story structure, scene design, character consistency, virtual camera language, sound, assembly, quality control, and the trade-offs behind common tool categories. It is written to be tool-agnostic on purpose, because model names change faster than the underlying craft.
The Three-Layer Workflow: Story, Shot, Signal
The cleanest way to organise an AI short film is three stacked layers. Each layer has its own document, its own review criteria, and its own failure modes.
Layer one — Story. The logline, the beat sheet, the character wants, and the turn. If this layer is weak, no amount of rendering quality saves the film. AI makes it easy to generate attractive footage; it does not make it easy to generate meaning.
Layer two — Shot. The scene list, the shot list, the coverage plan, the continuity rules. This is where you translate story beats into a sequence of discrete, renderable units. Each shot needs a beginning state, an end state, and a reason to exist.
Layer three — Signal. The prompt architecture: subject description, action, camera behaviour, lighting, lens, grain, palette, and negative constraints. Signal is the layer most creators obsess over and the layer that matters least if the two layers beneath it are vague.
A useful discipline is to freeze each layer before moving up. Do not write shot prompts while the beat sheet is still shifting. Do not design lighting while you are still deciding whether the protagonist has a sister. Freezing layers is what makes a three-minute film finishable in a week rather than a month.
A document stack that actually works
- A one-page treatment: logline, tone, three reference films, emotional arc.
- A beat sheet: eight to twelve beats with target durations.
- A scene bible: locations, time of day, weather, recurring props, colour logic.
- A character bible: age, build, wardrobe, hair, distinguishing features, voice notes.
- A shot list: numbered shots with duration, framing, movement, and dialogue.
- A prompt sheet: one row per shot with the full signal package.
Six documents, none longer than four pages. The point is not bureaucracy; it is that every regeneration decision has a written reference to check against.
Writing the Story Spine Before Touching a Tool
Generative models are excellent at producing the middle of things and terrible at deciding what the middle should mean. Write the spine on paper.
Logline and dramatic question
A short film needs one dramatic question and one visible answer. "Will the delivery rider open the envelope?" is a short film. "A delivery rider navigates a city and reflects on loneliness" is a mood reel. The first can be shot in twelve setups. The second has no natural end and will consume two hundred generations.
Write the logline in a single sentence with a clear subject, a goal, and an obstacle. Then write the question the audience will silently ask in the first twenty seconds. Every scene either advances that question or complicates it. Scenes that do neither get cut in pre-production, which is far cheaper than cutting them after rendering.
Beat sheet for a three-minute film
A reliable structure for a three-minute AI short looks like this:
- Cold open (0:00–0:15). A striking image that establishes world and tone. No exposition.
- Setup (0:15–0:45). Introduce the protagonist, the want, and the normal world.
- Disruption (0:45–1:10). The inciting incident that makes the want urgent.
- Escalation (1:10–1:50). Two or three attempts, each failing in a new way.
- Turn (1:50–2:20). The protagonist changes approach or learns the real cost.
- Climax (2:20–2:45). The decisive action, shot with the most visual emphasis.
- Release (2:45–3:00). A quiet final image that echoes the cold open.
This structure maps cleanly onto generation because each beat has a distinct visual register. Escalation beats need movement and changing light. The turn often needs stillness. The release needs a single held frame. If you can describe those registers in advance, you can choose models and settings per beat rather than applying one look to the whole film.
Dialogue that survives synthetic performance
AI voice and lip-sync still struggle with overlapping dialogue, heavy subtext, and rapid emotional pivots. In pre-production, write dialogue that plays to those constraints:
- Keep lines under twelve words.
- Give each character a distinct rhythm — one clipped, one meandering.
- Avoid lines that require an actor to cry and whisper simultaneously.
- Prefer off-screen dialogue where possible; it is far more forgiving and often more cinematic.
- Write silence deliberately. A beat of held silence between two lines reads as performance, whereas a gap created by a failed generation reads as a mistake.
Designing Scenes That Generative Models Can Actually Render
A scene is not just a location. It is a set of visual rules that must remain stable across every shot placed there. Writing those rules down is the single biggest quality improvement available to an AI filmmaker.
Location bible
For each location, record:
- Time of day and direction of key light.
- Dominant and accent colours, with hex values if you are precise.
- Floor and wall materials, since reflections are where models drift most.
- Two or three recognisable set dressing elements that must appear in every shot.
- A wide establishing reference you can attach as an image input.
If a scene takes place in a laundromat at dusk, the bible might specify: sodium-vapour practicals overhead, cold blue spill through the window on the left, cracked linoleum, one red plastic basket on the third machine. That single red basket becomes your continuity anchor. If it appears in three shots and disappears in the fourth, the audience will not consciously notice — but they will feel the scene come apart.
Complexity budgeting
Generative video fails predictably when a shot asks for too many simultaneous demands: multiple characters interacting, fast camera movement, detailed hand action, complex text, and dramatic lighting changes all at once. In pre-production, budget complexity per shot. Assign each shot a simple score and cap it.
A practical rule: no shot should contain more than two of the following — a speaking character, a moving camera, a hand interacting with an object, a crowd, visible text. Shots that exceed the cap get split into two shots or simplified in staging. This one discipline reduces reshoots more than any prompt trick.
Split for coverage, not just for safety
Because each shot is independently generated, coverage costs time but not much else. Design scenes with three registers: a wide that establishes geography, a medium that carries performance, and a close detail that carries emotion. Even if you only use two of the three, having a third gives the edit room to breathe when a generation comes back unusable.
Character Consistency: The Hardest Problem in AI Film
Audiences forgive imperfect sets. They do not forgive a face that changes shape between shots. Character consistency is therefore the pre-production task that most deserves obsessive attention.
Build a reference kit, not a reference image
A single portrait is not enough. Collect six to ten images of the character across angles: front, three-quarter left, three-quarter right, profile, and a full-body standing pose. Include one image in the primary costume and one in an alternate to understand how wardrobe alters the silhouette. Name the files systematically so you can attach the right references to the right shot without hunting.
Describe with anchors, not adjectives
Adjectives like "beautiful" or "tired" produce drift because every generation interprets them differently. Anchors are concrete and repeatable:
- Hair: shoulder-length, dark brown, tucked behind the left ear.
- Face: narrow jaw, straight brows, faint scar above the right eyebrow.
- Wardrobe: charcoal wool coat, unbuttoned, olive scarf.
- Distinguishing prop: silver ring on the right index finger.
Reuse this anchor block verbatim in every prompt. Never paraphrase it, never shorten it. Consistency comes from repetition, not from eloquence.
Multi-image fusion and when to use it
Many video tools accept more than one image input — a face reference plus a costume reference plus a lighting reference. Fusion is powerful but has a hierarchy problem: the model must decide which reference wins when they conflict. In pre-production, decide the hierarchy explicitly. A workable order is identity first, wardrobe second, lighting third. Then, in your shot notes, list references in that order every single time. Reordering inputs is a silent cause of character drift that many creators never diagnose.
Test before you commit
Before generating a single narrative shot, run a consistency test: the same character, three different framings, two different lighting conditions, one moving shot. If the face holds, proceed. If not, fix the anchors and references before producing anything you intend to keep. Two hours of testing saves a week of regeneration.
Shot Composition and Virtual Camera Language
AI video has made camera language more powerful and more fragile at the same time. You can now attempt a ninety-second unbroken take; you can also accidentally produce a shot with no readable geography.
Framing vocabulary to write into the shot list
- Extreme wide. Establishes scale and isolation. Use sparingly; detail degrades.
- Wide. Shows the body in space. Good for blocking and entrances.
- Medium. The workhorse for dialogue and reaction.
- Close. Carries emotion. Keep the background simple so the face holds.
- Insert. Hands, objects, screens. Cheap to generate, expensive to get right.
Write the framing word into the shot description itself. "Mara waits at the counter" becomes "Medium shot, Mara waits at the counter, camera static, chest height." The second version gives the model a decision it cannot get wrong.
Movement: the three safe moves
Slow push-in, slow lateral track, and static frame account for most successful AI shots. Fast pans, whip zooms, and handheld shake generate more artefacts than they generate energy. In pre-production, designate one signature movement for the film and use it at three or four emotional peaks. Repetition of a single movement creates style; a different movement in every shot creates noise.
Lighting as narrative
Because lighting drives so much of generation quality, treat it as a story element rather than a technical setting. Decide early whether the film's light gets warmer, colder, softer, or harsher as the story progresses. A protagonist moving from warm interiors to cold exteriors in eighty seconds communicates change without a single line of dialogue.
The 180-degree question
AI shots are frequently generated without a consistent spatial world, so continuity errors read less harshly than in live action. Still, keep a mental map per scene: where is the window, where is the door, which shoulder does the protagonist favour. Note it in the scene bible. When two consecutive shots contradict each other, the audience registers confusion even if they cannot name the cause.
Sound Design in an AI-Native Pipeline
Sound is where low-budget AI films most often give themselves away. The images may be coherent while the audio feels pasted on.
Plan audio in pre-production as three separate tracks: dialogue, ambience, and accents. Dialogue gets recorded or synthesised per line and checked against lip movement. Ambience gets built per location — a room tone for the laundromat, a street bed for the exterior — and must be contiguous across cuts within the same scene. Accents are single events: a door latch, a phone buzz, a chair scrape.
Practical rules
- Build one ambience bed per location, not per shot. Reuse it; do not regenerate.
- Cut picture to sound where possible. A hard cut landing on a door slam feels intentional; a cut landing on nothing feels amateurish.
- Keep music out until the picture is locked. Temp music disguises weak structure and will make you keep scenes that should be cut.
- Normalise dialogue to a consistent level before mixing, so the audience is not riding the volume.
- Silence is a tool. Two seconds of clean silence before a turn is more effective than a swell.
Voice synthesis has improved enough that a single performer can plausibly voice two characters with pitch and rhythm changes, but only if the writing differentiates them. If both characters speak in identical sentence structures, the audience will hear one voice regardless of the pitch.
Assembly, Quality Control, and Common Mistakes
Once shots exist, the edit becomes an exercise in brutal selection. Build a rough assembly with placeholder cards for missing shots, then grade every shot against four criteria.
The four-point check
- Continuity. Does the face, wardrobe, and location match the scene bible?
- Motion. Is the movement physically plausible, or does it warp mid-shot?
- Purpose. Does the shot advance the beat, or is it merely attractive?
- Duration. Is it as short as it can be while still readable?
Shots that fail two or more criteria get regenerated. Shots that fail only duration get trimmed. Shots that are beautiful but purposeless get cut, and this is the hardest discipline in the entire workflow.
Mistakes that cost the most time
- Prompt sprawl. Rewriting the entire prompt when one variable changed, then losing the version that worked. Version your prompts like code.
- Fixing in post. Attempting to salvage a drifting face with stabilisation and masks. Regenerate instead; masking a moving face in a synthetic shot is rarely worth it.
- Unlimited coverage. Generating forty options for one shot and losing the thread of the film. Cap options at three per shot.
- No duration target. Letting shots run to the model's default length. Every shot in your list should have a target duration before you generate it.
- Music-led editing. Cutting to a track instead of to story beats. It feels good in the rough cut and hollow in the finished film.
Tooling Landscape and How to Choose
Rather than chasing a single best model, choose tools per task. Categories matter more than brands.
Text-to-video and image-to-video engines differ mainly in motion realism, prompt adherence, and maximum clip length. Test three with the same prompt on your own footage before committing. Some excel at landscapes and struggle with faces; others do the reverse.
Image generators remain the best way to lock a look before animating it. Generate a still first, approve it, then animate. Starting from an approved still raises the floor of every subsequent render.
Consistency and reference tools — face-swap utilities, identity adapters, and multi-reference workflows — are worth the setup cost on any film with a recurring protagonist.
Voice and audio tools should be selected for control, not novelty. You want per-line generation, pitch and pacing control, and a clean export format.
NLE and finishing tools matter more than people expect. A capable editor with good trimming, proxy handling, and audio metering will save more time than a faster generator. Colour management matters too; mixing shots from different models requires matching contrast, saturation, and grain before the film looks like one film.
A practical selection rule: pick one engine for character work, one for environments, and one for inserts, then resist switching mid-production. Style coherence comes from a stable pipeline far more than from any single model's quality.
FAQ
How long should an AI short film be?
Ninety seconds to three minutes is the practical sweet spot. Long enough for a real turn, short enough that consistency problems do not compound. If you are new to the pipeline, aim for sixty seconds and finish it.
Do I need a full storyboard?
No, but you need a shot list with framing, duration, and a one-line action description. Storyboards help with complex blocking; simple annotated stills are usually enough.
What is the most common reason AI short films fail?
Weak story spines and character drift, in that order. Beautiful footage of a story that never turns produces an empty film, and a protagonist whose face changes every shot produces confusion that no score can fix.
How many generations should I expect per finished shot?
With a locked reference kit and a written anchor block, two to four is realistic. Without them, ten or more is normal, and the extra attempts rarely converge on the same character.
Should I generate video directly from text or from an image?
From an image, whenever the shot contains a recurring character or a specific location. Text-to-video is best for establishing shots, textures, and transitions where continuity is not at stake.
How do I keep lighting consistent across a scene?
Specify light direction, colour temperature, and source in the scene bible, then repeat that language verbatim in every prompt for that location. Changing the wording changes the light.
Can one person make a short film like this?
Yes, and many do. The realistic division of labour is writing, generation, and editing, each getting a dedicated block of time. Trying to do all three simultaneously is what makes solo production feel impossible.
What should I do first if I only have a weekend?
Write the logline and beat sheet on paper, build one character reference kit, and generate a single thirty-second proof of concept containing three shots. If the character holds across those three shots, you have a film. If not, you have saved yourself a weekend of unusable footage.



