What an AI Short Film Pipeline Actually Looks Like
Making a short film with generative video is not the same as making one with a camera, and the difference is not just technical. When you shoot on set, your job is to capture what happens in front of the lens. When you generate, your job is to describe what should happen and then select the best interpretation of that description. You are not filming reality â you are directing probabilities.
That shift changes everything about the workflow. There is no call sheet, no lighting crew, no reshoot budget. Instead you get three phases that repeat in a tight loop:
- Pre-production â story, beat sheet, shot list, and a style bible that keeps every generated clip looking like it belongs to the same film.
- Generation â turning each shot into a prompt, picking the right model for that shot, and generating enough variations to have real choices in the edit.
- Post-production â cutting, sound design, colour, and export.
Most beginners skip straight to generation, spend four hours producing beautiful disconnected clips, and then discover they cannot cut them together into anything coherent. The short film you want lives or dies in the first phase, and the good news is that pre-production for a two-minute AI film takes about thirty to forty-five minutes if you do it properly. That half hour routinely saves three hours of regenerating shots that never fit.
A realistic time budget for a first 60-second film looks like this: 40 minutes of planning, 2 to 3 hours of generating and re-rolling, 1 to 2 hours of editing and sound. It is entirely possible to finish in an afternoon, but only if you resist the urge to start generating before you know what you are making.
Pre-Production: The Work That Decides Everything
Write the logline and beat sheet first
A logline is one sentence: who wants what, what stands in the way, and what is at stake. "A lighthouse keeper discovers the beam is answering a signal from the sea." That is enough to generate from. Everything after that is expansion.
Next, break the logline into six to eight beats. Each beat is a change in the situation, not a mood. A beat sheet for a 60-second film might read: opening image, inciting detail, first attempt, complication, turn, climax, final image. When you generate shots, every shot will map to a beat. If you cannot say which beat a shot serves, cut it.
Build a shot list with four columns
Your shot list is the single most useful document in the whole project. Four columns are enough:
- Shot number and duration â aim for 2 to 4 seconds per shot in a first film. Long generated shots have more room to drift, morph, or reveal artifacts.
- Framing â wide, medium, close-up, insert, over-the-shoulder.
- Action â what physically changes in the frame.
- Mood and light â time of day, colour temperature, weather, contrast.
A 60-second film at a 3-second average is 20 shots. That feels like a lot until you realise how much coverage you need for a simple scene: an establishing wide, a medium of each character, a reaction close-up, an insert of the object that matters, and a final wide. Five shots, fifteen seconds, one location. Multiply that across four locations and you are already at 20 shots without any padding.
Create a style bible
A style bible is a short list of fixed decisions you repeat in every prompt. Write it down and never improvise on set:
- Palette â for example, teal shadows, sodium-orange highlights, desaturated greens.
- Lens character â 35mm spherical with subtle halation, or anamorphic with horizontal flares.
- Film stock or render look â grainy 16mm, clean digital, painterly animation.
- Lighting logic â soft overcast, hard single-source, practical neon.
- Camera behaviour â locked-off and composed, or handheld and reactive.
Consistency comes from repetition of these phrases far more than from any single model setting. Two clips generated by different tools will still feel like the same film if the palette, lens, and lighting language match.
Choosing a Generation Model Shot by Shot
There is no single best video model, and treating this as a one-tool decision is the most common strategic error beginners make. Models differ along axes that matter differently per shot.
The criteria that actually matter
- Clip length â some tools cap out at a few seconds; others allow sustained takes. Wide establishing shots usually tolerate shorter clips; dialogue scenes need longer.
- Motion realism â how convincingly bodies move, walk, and interact. Weak here means you should avoid full-body action entirely.
- Prompt adherence â how literally the output matches your description. High adherence is essential for continuity-critical shots.
- Image-to-video support â whether you can drive the shot from a still you already approved.
- Native audio â some models generate ambience or dialogue; most do not. Plan for external sound if yours does not.
- Style bias â models have house looks. Some lean photoreal and cinematic, others lean stylised and animated. Fighting a model's bias wastes iterations.
A practical matching approach
Describe each shot in one line, then tag it with the property it needs most. A misty landscape at sunrise needs atmosphere and slow camera motion. A close-up of a hand opening a letter needs subject detail and stability. A chase needs motion coherence and fast cutting. Then assign tools: use the stylistically strongest option for hero shots, and a faster, cheaper-feeling option for connective shots like inserts and establishing plates that will be on screen for under two seconds.
When to switch mid-project
Switch tools when a shot has failed three times for the same reason. If every attempt produces the same artifact â a face melting during a turn, hands fusing, a street set that shifts geometry â the model is the problem, not the prompt. Three failed attempts for the same identifiable cause is your signal to change approach, either by changing tool, changing the shot type (from a full turn to a profile glance), or changing the framing (wider so detail is less scrutinised).
Prompting for Motion and Camera Language
Anatomy of a strong video prompt
A reliable structure, in order:
- Subject â who or what, with two or three identifying details.
- Action â a single continuous movement. One verb phrase per shot.
- Camera â the movement and the lens.
- Light and atmosphere â the source, direction, and quality of light.
- Style anchor â the phrases from your style bible.
Example: "A weathered lighthouse keeper in an oilskin coat lifts a brass lantern; slow dolly in on a 35mm lens; cold blue dusk light from the west with a faint amber glow from the lantern; fine grain, teal shadows, muted greens."
Notice that the action is one continuous physical movement. Prompts that ask for two sequential actions ("she opens the door and then runs down the stairs") usually produce a compromise that does neither well. Split it into two shots.
Camera verbs that work reliably
The camera vocabulary models understand best is physical and directional:
- slow dolly in / slow push out
- orbit around the subject, left to right
- handheld follow behind the subject
- crane rise revealing the landscape
- lateral tracking shot, camera remains level
- locked-off static frame, no camera movement
The last one is underrated. A locked-off shot gives the model one less variable to corrupt, and static frames cut together beautifully when the subject moves within them.
Handling artifacts
Expect these and plan around them: hands morphing, faces drifting between identities, background geometry warping, limbs duplicating, and textures boiling. Mitigations, in order of effectiveness: shorten the clip; reframe tighter or wider to remove the problem area; convert to a static camera; drive the shot from a reference still instead of text; and finally, cover it in the edit with an insert or a cutaway.
Character and Location Consistency
Start from a reference still
The most reliable consistency method is to approve a still image of your character first, then use image-to-video for every shot that character appears in. This gives the model a fixed anchor for face, wardrobe, and proportions, and it removes most of the guesswork from text-only prompting. Generate the still, refine it, keep it, and reuse it relentlessly.
Anchor with props and wardrobe
Identity in film is carried by costume and props as much as by faces. A red scarf, a specific jacket, a scar, a pocket watch â these give the audience continuity markers and give your prompts concrete language to repeat. Lock two or three anchors per character and repeat those exact phrases in every prompt for that character.
Handle scene changes without breaking the film
When a character moves to a new location, generate a clean establishing shot before any character shots there. The establishing shot sets the light and geography, and you can then drive character shots from stills that use that same lighting description. If the film jumps time or place, a deliberate hard cut in lighting logic reads as intentional; accidental drift between two shots in the same scene reads as a mistake.
Assembling the Cut
Generate coverage, not single perfect takes
Beginners generate one clip per shot and hope. Better: generate three to five variations per shot, then select. The cost of a re-roll is seconds; the cost of a bad take in a finished edit is your whole film's credibility. Treat generation as shooting coverage.
Cut on motion and match continuity
AI footage cuts best when the cut lands on movement. If a subject is walking left, cut as they cross the centre of the frame. If a camera is dollying in, cut at the moment of maximum forward energy. Match screen direction, eyeline, and light between adjacent shots, and keep a consistent grade across the timeline. Slight speed changes â 90% or 110% â often hide small motion mismatches and are invisible to an audience.
Fix warping with three moves
When a clip has visible drift or morphing near the end, the answer is usually one of three things: trim before the artifact appears; ramp the speed so the artifact passes faster; or cover the moment with an insert or reaction shot. Never leave a morph on screen for a full second hoping nobody notices â they will.
Sound Design
Sound is the fastest way to make generated footage feel like a real film, and it is the step beginners skip most often.
Decide on dialogue early
If your film has dialogue, generate voice separately and cut the picture to the audio, not the other way around. Lip-sync from generative video is still fragile, so shoot dialogue in profile, in medium shots, from behind, or cut away to the listener during lines. Alternatively, embrace the silent-film approach: no dialogue, strong visuals, a music-led mix, and title cards for essential information.
Layer three audio bands
- Ambience â a continuous bed for each location: sea, city hum, forest, room tone.
- Foley â specific sounds tied to on-screen action: footsteps, cloth, a latch, a match striking.
- Music â a single cue with one clear arc. Do not use three themes in sixty seconds.
Mix targets that work
Keep music under dialogue, cut ambience when the location changes, and add two seconds of room tone at the head and tail so the film does not start and end in dead silence. A simple loudness-normalised export around â14 LUFS for streaming platforms is a safe default.
A Worked Example: 60-Second Short in One Afternoon
Here is a real sequence of work for a 60-second atmospheric short about a diver discovering something below a frozen lake.
0:00â0:40 â Planning. Logline written. Eight beats. Twelve shots listed: three establishing plates of the frozen lake, two shots of the diver gearing up, four underwater shots, two reaction close-ups, one final wide. Style bible fixed: cold cyan palette, 35mm, hard top-light through ice, fine grain, mostly locked-off with two slow pushes.
0:40â1:10 â Stills. Character still of the diver developed and approved. Three environment stills generated for the lake surface, the entry point, and the underwater cavern.
1:10â3:30 â Generation. Each shot driven from a still where possible, three variations each. Roughly twenty of thirty-six clips are usable. Four shots require a switch to a different model after three identical failures.
3:30â4:15 â Assembly. Rough cut in order, cutting on motion. Two shots trimmed to 1.5 seconds to hide morphing. One shot replaced by an insert.
4:15â5:00 â Sound. Ambience bed for surface and underwater, foley for regulator breathing and ice cracking, one music cue.
5:00â5:30 â QC and export. Check below, then render.
Quality control checklist before you export: no visible morphs in any frame; consistent character wardrobe across every appearance; screen direction consistent; no shot longer than four seconds unless deliberately held; audio peaks below clipping; no dead air at head or tail; grade consistent start to finish; and the final image lands emotionally rather than just ending.
Common Mistakes and Troubleshooting
Too many shots, too little runtime
Twenty shots in sixty seconds means three seconds each, which is fast cutting. Beginners often list forty shots, then cannot fit them without the film feeling like a trailer. Count your shots against your runtime before generating.
Ignoring the first eight frames
Many models produce a slight settle in their opening frames. Start your edit a few frames in, or generate a longer clip and trim forward. This single habit removes a lot of amateur feel.
No continuity document
Keep a running note of each character's wardrobe, each location's lighting, and each prop's position. When you are on shot eighteen at hour three, you will not remember whether the scarf was on the left or right shoulder. The note will.
Chasing a single perfect take
If a shot resists ten attempts, the shot is wrong, not the prompt. Reframe it, shorten it, or cut it entirely. Films are made in the edit, and a missing shot is often an improvement.
Neglecting the grade
Applying one consistent colour treatment across every clip is the cheapest, fastest way to make footage from different tools feel like one film. Do it before you judge the cut.
FAQ
How long should my first AI short film be?
Sixty to ninety seconds. It is long enough to tell a real story and short enough to finish in one or two sessions. Feature-length ambitions should wait until you have completed three short pieces.
Do I need editing software?
Any timeline-based editor works, including free options. You need trimming, speed changes, a colour grade, and at least two audio tracks. Nothing more advanced than that is required for a first film.
Can I make a film in one tool?
You can, and it is a reasonable way to start because it removes the friction of moving files around. As your ambition grows, using different models for different shot types becomes normal and improves results.
How do I stop characters from changing between shots?
Approve a still first, then drive every shot from it. Repeat exact wardrobe and feature phrases in each prompt, and avoid shots that turn a face through large angles.
Is text-to-video or image-to-video better for beginners?
Image-to-video is more controllable and produces far fewer surprises. Text-to-video is useful for establishing shots and abstract transitions where nothing specific needs to persist.
How many variations per shot should I generate?
Three is the practical minimum, five if the shot is a hero moment. Select in a contact-sheet view rather than watching clips in full, which saves a surprising amount of time.
What kills an AI short film fastest?
Inconsistent lighting between shots in the same scene, visible morphing held on screen, and a music bed that never changes. All three are fixable in post, and all three are why planning beats generating.
Do I need permissions or clearances?
Check the licensing terms of every tool and asset you use, including music, and confirm what commercial use is permitted before publishing. Keep a simple record of the assets in your project so you can answer questions later.




