Why text-to-video changed the production pipeline
A decade ago, a sixty-second brand film meant a crew, a location, a lighting kit, talent releases, and a week in the edit. Text-to-video generation collapsed most of that into a browser tab. You describe a shot in language, the model renders motion, and you iterate on wording instead of logistics. The camera department becomes a paragraph.
The real change is not speed, though the speed is considerable. It is the feedback loop. A director can test three visual approaches before lunch, keep the one that works, and discard the rest without reshooting anything. Cost per attempt drops far enough that exploration becomes normal rather than exceptional.
That shift moves the bottleneck. Production capacity is no longer the constraint; decision quality is. Teams that produce strong AI video are rarely the ones with the most exotic models. They are the ones with a clear shot list, a consistent vocabulary for describing images, and a disciplined review process that kills weak clips early.
This guide lays out a practical pipeline for turning a written script into a finished, coherent video using generative models: how to plan shots, how to pick a model per shot, how to prompt for motion, how to keep characters and locations stable across cuts, and how to review before export.
The four layers of an AI video workflow
Almost every successful AI video project passes through the same four layers. Skipping one usually shows up later as rework.
Layer one: script and shot list
Write the script first, in plain prose, the way you would for a human crew. Then translate it into shots. A useful rule: roughly 90 to 110 words of narration maps to eight to twelve shots for a one-minute piece. Each shot entry should record five things: subject, action, camera behaviour, environment, and duration. A shot that says 'she opens the letter, close on hands, warm window light, three seconds' is ready to generate. A shot that says 'emotional moment' is not.
Layer two: generation
Generate three to five candidate clips per shot rather than one perfect take. Variation is cheap; a missing option is expensive. Save every candidate in a folder named after the shot number, not after the prompt, so the edit stays organised when prompts change.
Layer three: assembly
Cut to rhythm before you cut to beauty. Lay the candidates on a timeline, find the pace, and only then replace weak shots with better generations. Many projects fail because the editor falls in love with a gorgeous clip that breaks the rhythm of the sequence.
Layer four: polish
Colour match, sound design, music, captions, and loudness normalisation. This is where amateur AI video stops looking like a demo. Even a light grade and a proper audio bed lifts generated footage significantly.
Choosing the right model for each shot
No single model wins every category. A practical approach is to assign models by shot type, then verify with a short test batch before committing to a full sequence.
| Shot type | What matters most | Practical approach |
|---|---|---|
| Realistic human close-up | Facial stability, skin texture | Test two or three models, pick one, reuse it for all dialogue-adjacent shots |
| Wide establishing shot | Depth, atmosphere, camera motion | Favour engines strong on landscape and drone-style moves |
| Product macro | Fine detail, controlled lighting | Image-to-video from a still keyframe gives the most control |
| Stylised or animated | Consistent art direction | Prefer models that hold a visual style across prompts |
| Complex action | Physics plausibility | Keep cuts short; long action beats degrade quickly |
Three decision criteria matter more than leaderboard rankings.
First, motion coherence over still quality. A model that produces beautiful frames with warping limbs is less useful than one with slightly softer detail and stable motion, because motion errors are almost impossible to fix in post.
Second, duration per generation. If a model reliably produces five usable seconds, plan five-second shots. Fighting for a ten-second take usually costs more time than adding a cut.
Third, style flexibility. Some models have a strong house look. That is an advantage for a single project and a liability across a client roster. Test the same prompt across candidates and compare skin tones, contrast, and colour bias before you build a house style around one engine.
A useful habit: keep a personal model notes file. Record which model you used for which shot type, what the prompt looked like, and whether the result needed heavy repair. After twenty projects, that file is worth more than any comparison chart.
Writing prompts that survive the jump from text to motion
The five-part skeleton
Most reliable prompts follow a consistent order: subject, action, camera, environment, style. For example: 'A ceramicist in her sixties, hands shaping a bowl on a spinning wheel, slow push-in from a low angle, dusty sunlit studio with shelves of unfinished pots, warm documentary realism, shallow depth of field.'
The order matters less than the completeness. What breaks prompts is omission, usually camera or environment, and then the model invents something you did not want.
Words and ideas that cause trouble
- Negations. 'No crowd' often produces a crowd. Describe the desired state instead: 'empty street at dawn.'
- Abstract emotions. 'She feels betrayed' produces a generic expression. Describe behaviour: 'she looks away, jaw tight, then closes the notebook.'
- Multiple simultaneous actions. 'He runs, types, and laughs' confuses motion planning. Keep one primary action per clip.
- Dialogue. Most models do not render believable speech. Generate the performance, then dub the line.
Iterate one variable at a time
When a clip fails, change exactly one element, camera first, then lighting, then wardrobe, and regenerate. Changing three at once tells you the result improved but not why, and you cannot repeat the win. Keep the prompt text in a notes column next to the clip so you can trace which phrasing worked.
Maintaining visual consistency across shots
Consistency is the hardest part of AI video, and it is where most projects visibly break down. Characters change faces between cuts, rooms rearrange themselves, and colour drifts from warm to clinical.
Lock a reference frame
Generate one strong image of each character and each key location. Approve it before generating motion. From then on, use that frame as the starting image for image-to-video, or describe it in identical words at the start of every prompt. Written consistency compounds: if the wardrobe line is always 'olive wool coat, cream scarf', the model has less room to improvise.
Fix the visual grammar
Decide early on lens language and lighting logic. If the film uses a 35 mm feel with soft window light in every interior, say so in every interior prompt. Colour drift usually comes from inconsistent style words rather than from the model itself.
Reuse seeds where available
Many generation tools let you keep a seed value for repeatability. Reusing a seed with a lightly edited prompt often preserves background structure and lighting while changing the action. It is not a guarantee, but it raises your hit rate noticeably.
Build a character sheet
One page per recurring character: age, build, hair, wardrobe, two reference stills, and the exact prompt sentence used to describe them. Share it with anyone else prompting on the project. This single document prevents most continuity failures.
Sound, voice, and pacing
Generated video is silent, and silence is what makes it feel artificial. Three layers fix that quickly.
Voice. Write for the ear, not the page. Short sentences, concrete nouns, no clauses that require a breath you cannot hear. Record a scratch track yourself before generating anything, so shot lengths match the narration rather than the other way around. If you use synthetic voice, keep it slightly slower than feels natural on the page; listeners need a moment to absorb a new visual.
Ambience. A room tone under every scene, even a quiet one. Wind, distant traffic, a refrigerator hum. Ambience is the cheapest way to make generated footage feel filmed.
Music. Choose the track before final assembly, not after. Tempo dictates cut points, and cutting to a beat is the fastest route to a polished feel.
Rhythm. Vary shot length deliberately: two seconds, two, four, one, three. Uniform shot lengths, especially uniform five-second generations, read as a slideshow. Sound design also covers motion seams. A whoosh, a click, or a fabric rustle can hide a transition that is not perfectly smooth.
Finally, check loudness. Aim for a consistent integrated level across the whole piece so viewers do not reach for the volume control between sections.
A worked example: a 45-second teaser
Suppose the brief is a teaser for a small-batch coffee roaster. The script is 80 words of narration. The shot list has nine shots.
- Dawn exterior of the roastery, slow drift right, 4 seconds.
- Close on hands opening a burlap sack, 2 seconds.
- Beans falling into the drum, macro, 2 seconds.
- Roastmaster watching the drum, medium close-up, 3 seconds.
- Steam and smoke in a shaft of light, wide, 3 seconds.
- Beans tumbling, overhead, 3 seconds.
- Cup being filled, slow motion, 4 seconds.
- First sip, close on face, 3 seconds.
- Exterior at dusk, lights on, pull back, 5 seconds.
Assign models by shot type. Macro product shots go to an image-to-video workflow seeded with stills of actual beans, because real texture matters for a food product. Human close-ups use a single model throughout so the roastmaster's face stays recognisable. The two exterior establishing shots use whichever engine handles atmosphere and camera drift best.
Generate four candidates for each shot in a single batch. Expect roughly half to be usable and a quarter to be good. Assemble a rough cut at 45 seconds, then replace the weakest three shots. Add ambience, a music bed at a tempo around 90 BPM, and narration. Grade everything to one warm look so the exteriors and interiors feel like the same film.
The total active work is a few hours. The same project shot conventionally would need a location day, a food stylist, and a permit.
Common mistakes and how to avoid them
- Prompting before planning. Generating clips without a shot list produces attractive orphans that never cut together.
- Chasing duration. Long generations degrade. Cut more often.
- Ignoring frame one. If the first frame is wrong, the motion will be wrong too. Fix the still before the motion.
- Mixing too many models in one sequence. Each engine has a colour bias; a sequence stitched from five engines looks stitched.
- Overwriting prompts. Adding adjectives to fix a problem usually adds noise. Remove words before you add them.
- Neglecting audio. Silent drafts get approved and then fall apart in review.
- Accepting warped hands. Viewers notice. Regenerate rather than crop.
- Skipping the small-screen check. Watch the cut on a phone before export. Fine detail that reads on a monitor often disappears.
- No naming convention. Unlabelled clips turn a two-hour edit into a two-day hunt.
- Never deleting anything. Keep your archive, but clear the working timeline. Clutter slows decisions.
Review checklist before you export
- Continuity: does each recurring character keep the same face, wardrobe, and apparent age?
- Geography: do locations make spatial sense from shot to shot?
- Motion: any warping limbs, melting objects, or flickering textures?
- Colour: one consistent grade across all shots?
- Audio: ambience under every scene, music balanced, narration intelligible?
- Loudness: consistent level from first frame to last?
- Captions: accurate, readable at phone size, no clipping?
- Ratio and safe areas: correct aspect ratio, key text away from edges?
- Licensing: do you have rights to every still, track, and voice you used?
- File naming: project, version, and date-stamped exports in one folder?
Run the list twice: once with sound off, once with picture hidden. Each pass catches different problems.
FAQ
How long should an AI-generated shot be?
Two to five seconds covers most needs. If you need a longer continuous take, generate two clips from the same reference frame and cut between them, or use the second as a continuation. Long single generations tend to drift.
Can I use AI video for client work?
Usually yes, with two caveats: check the licence terms of each model you use, and disclose synthetic imagery where your client's industry requires it. Food, fashion, and finance clients often have stricter internal rules than the law does.
Do I need a powerful computer?
Not necessarily. Most generation happens in the cloud. Local work is mostly editing, where any modern laptop with a mid-range GPU handles 1080p comfortably. If you deliver 4K regularly, plan for more storage rather than more compute.
How do I keep a character's face consistent?
Generate one approved still, use it as the starting frame for every shot featuring that character, and keep the descriptive sentence identical across prompts. Reusing a seed helps. Nothing guarantees perfection, so favour shots that show the face briefly rather than holding a long close-up.
Is a storyboard still useful?
More useful than ever. It forces you to decide what each shot communicates and gives you a checklist for generation. Rough stick figures are enough.
What is the fastest way to improve?
Rebuild a thirty-second scene from a film you admire, shot for shot. You will learn more about pacing, coverage, and light from one exercise than from a month of random prompting. Keep both versions side by side and write down the three biggest differences.


