Why a Workflow Beats a Single Prompt
Generative video has reached the point where almost anyone can produce a striking five-second clip. What separates a finished video from a folder of attractive fragments is not the model you use. It is the process wrapped around it. A repeatable workflow gives you three things a clever prompt cannot: predictability, continuity, and a way to diagnose failure. When a shot comes back wrong, a workflow tells you whether the problem was the reference image, the motion instruction, the aspect ratio, or simply an unlucky roll.
There is also a compounding cost to skipping process. Without a workflow you regenerate rather than repair, and every regeneration resets continuity you had already solved. Five unplanned retries on a character shot can cost more time than the entire script and storyboard phase combined. The goal is not to eliminate iteration, it is to make iteration cheap, directed, and reversible.
Treat the model as one station on an assembly line. The line starts with a brief and ends with a delivered file, and every station has an input specification and a quality gate. Build that line once and you can swap models in and out as new ones appear without rebuilding your entire production habit. When a new model arrives, you test it against three known shots and decide whether it replaces a station or simply adds an option.
Stage 1: Define the Brief Before You Open Any Model
Format, runtime, and aspect ratio
Decide delivery specifications before generating a single frame. Vertical 9:16 for short-form social, 16:9 for landscape storytelling, 1:1 or 4:5 for feed placements. Runtime matters more than most people expect: a 30-second piece built from eight shots averages under four seconds per shot, which means each clip must carry a complete visual idea. If you sketch a 30-second script with 20 shots, you have written a three-second-per-shot edit that will feel frantic no matter how beautiful the individual clips are.
Aspect ratio decisions made late are expensive. Generating in 16:9 and then cropping to 9:16 throws away more than half your pixels. Generate in the ratio you will deliver, or shoot for the widest ratio and compose with safe margins if you know you need multiple cutdowns.
Tone and treatment
Write two or three sentences describing the feel of the piece: naturalistic handheld documentary, clean studio product, stylised neo-noir. This treatment paragraph becomes the reference for every style phrase in every prompt. Without it, shot five drifts into a different genre than shot one, and no amount of colour grading fully repairs that.
The shot list
Write each shot as a single sentence with a subject, an action, and a camera behaviour. A woman in a yellow raincoat steps off a tram, camera tracks left at walking pace is a shootable instruction. A moody city vibe is not. Number your shots, note which ones need the same character or location, and mark which are essential versus optional. Optional shots are your buffer when a model refuses to cooperate with a difficult action.
The reference pack
Collect still images for every recurring element: face, wardrobe, vehicle, location, product. High resolution, neutral lighting, one subject per image, no overlapping clutter. These references will do more for consistency than any adjective you can put in a prompt.
Stage 2: Choose the Model That Matches the Shot
Text-to-video, image-to-video, and video-to-video
Text-to-video is for establishing shots, abstract transitions, and anything where you do not yet have a reference. Image-to-video is for shots anchored to a character or product, where you supply the first frame and describe motion only. Video-to-video and motion-transfer tools are for restyling existing footage or borrowing a movement from a reference clip. Skilled editors mix all three in one timeline: an establishing aerial from text, character coverage from stills, and a stylised insert from existing footage.
What to test when comparing models
Ignore leaderboard scores and run your own three-shot test: one portrait with subtle facial motion, one wide landscape with camera movement, and one shot of a hand interacting with an object. Score each on motion realism, prompt adherence, temporal stability, and how faithfully it preserves a supplied reference. The model that wins on hands and faces is usually the one worth building around, even if another produces prettier landscapes.
Run this test again whenever a model gets a significant update. Assumptions about which tool handles which shot age badly, and a model that struggled with hands six months ago may now be your best option for them.
Match the model to the moment
Fast, inexpensive generation is right for exploration. Higher-quality, slower generation is right for hero shots. Planning around a single model for everything is the most common cause of blown schedules, because you either overspend quality on shots nobody will notice or underspend it on the shots that carry the piece.
Build a hybrid chain
Most strong pipelines use three or four tools rather than one. A typical chain: text-to-video for establishing shots, image-to-video for character coverage, a dedicated upscaler for accepted takes, and a separate audio tool for voice and score. Document which tool handles which shot type so a collaborator can reproduce your results.
Stage 3: Prompt Structure That Survives Generation
The five-part shot prompt
Subject, action, environment, camera, and style. Write them in that order. A middle-aged fisherman hauls a net over the gunwale on a foggy wooden boat at dawn, slow handheld push in from behind, muted documentary colour with soft grain. This order mirrors how most models weight language and keeps you from burying the important part at the end of a long sentence.
Camera and lighting vocabulary
Use terms the model has seen in captions: dolly in, tracking shot, crane up, static locked-off, handheld, drone orbit, shallow depth of field, golden hour, overcast diffused light, practical neon, hard key light from the left. Vague words like cinematic carry little weight on their own. Pair them with a concrete light source and a lens behaviour, and the output stops looking generic.
Negative instructions and restraint
Describe what you want rather than listing what you do not. If a shot keeps adding unwanted elements such as crowds, extra limbs, or text overlays, simplify the prompt instead of stacking negations. Fewer nouns means fewer things that can go wrong. Keep prompts under roughly 60 to 80 words for motion-heavy shots, and expand only when a model demonstrably ignores detail you have already included.
Dialogue and on-screen text
Treat spoken lines and visible text as post-production problems, not generation problems. Lip-synced generation has improved but remains the least reliable part of the chain. Generate clean plates, then add dialogue with a separate voice tool and add on-screen text in the edit where you control font, timing, and safe margins.
Stage 4: Hold Consistency Across Shots
Reference frames and multi-image conditioning
Give the model the same face reference for every shot featuring that character, and the same location reference for every shot in that space. When a model supports multiple image inputs, use them deliberately: one for identity, one for wardrobe, one for environment. Blending five references indiscriminately produces a person who looks like nobody.
Continuity notes
Maintain a one-page continuity sheet listing hair, clothing, props, time of day, and colour temperature for each scene. It sounds like paperwork. In practice it is the difference between a sequence that reads as a film and one that reads as a montage of unrelated clips. Update it as decisions change rather than reconstructing it at the end.
Lock hero frames early
Once you have a first frame you love, export it and reuse it rather than regenerating. Every regeneration is a new roll of the dice, and consistency is far cheaper to preserve than to repair.
Camera continuity
Continuity is not only about faces. If shot three cuts from a leftward track, reversing to a rightward track in shot four can read as a jump even when the subject matches. Note screen direction, eyeline, and the side of frame each character occupies.
Stage 5: Generate, Review, and Accept Takes
Keep a take log
For each shot, record the prompt version, reference images, seed if the tool exposes one, and a one-word verdict. After twenty generations you will not remember which combination worked, and the log turns a lucky accident into a repeatable setting you can hand to a collaborator.
Read failures precisely
| Symptom | Likely cause | First fix |
|---|---|---|
| Melting faces, identity drift | Weak or conflicting reference | Use a single clean frontal reference |
| Jittery motion, warping edges | Too much motion per second | Shorten the clip, slow the action, steady the camera |
| Prompt largely ignored | Overlong or contradictory prompt | Cut to subject plus action, move style to the tail |
| Unwanted text or logos | Model bias toward signage | Specify clean surfaces, reframe the shot |
| Colour shifts between shots | Inconsistent lighting description | Standardise light and correct in the grade |
Batch by shot type
Generate all portraits together, then all landscapes, then all inserts. Switching between prompt styles and reference sets is slower than it feels, and batching also makes inconsistencies easier to spot because similar shots sit side by side.
Accept good enough on time
Aim for most shots landing within one or two attempts and reserve your patience for the hero shots. Perfectionism on shot 14 of 20 is how deadlines die, especially when the audience will see that shot for 1.5 seconds.
Stage 6: Post-Production in an AI-Native Timeline
Upscale, interpolate, stabilise
Generated clips often arrive at modest resolution with slight temporal inconsistency. Upscale before you grade, not after. Frame interpolation smooths slow motion but can introduce ghosting on fast action, so test before committing to it across a sequence. Stabilisation should be light; aggressive stabilisation applied to a shot with deliberate handheld movement destroys the intent you prompted for.
Sound carries the illusion
Audiences forgive imperfect frames far more readily than bad audio. Lay a bed of ambience for every location, add foley for visible actions such as footsteps, fabric, doors, and water, and use music to cover cuts. Generate or record voice separately and align it to the edit rather than trying to make a clip's length fit a line of dialogue.
Colour and finishing
Do a single grade across the whole sequence. Match black levels and white balance between shots before adding any look. A consistent, restrained grade makes generated footage feel noticeably more professional than any individual shot improvement ever will.
Export and delivery settings
Check delivery specifications before exporting: codec, bitrate, frame rate, loudness target, caption format. Rendering a second time because the frame rate was wrong is the most avoidable delay in the whole pipeline.
Quality control checklist
- Play the full timeline at normal speed without pausing
- Mute the audio and watch: do the visuals still tell the story?
- Watch on a phone at arm's length to approximate the real viewing context
- Inspect the first and last two seconds of every clip for artefacts
- Verify captions, safe margins, and loudness
Common Mistakes and How to Avoid Them
Writing a script instead of a shot list is the first mistake. Video models render moments, not scenes, and a paragraph of narration will not map cleanly onto a sequence of clips. Generating at the final aspect ratio before you understand the edit is the second: crop decisions made late cost resolution you cannot recover.
Chasing a single perfect model rather than a working chain is the third mistake. No model is best at portraits, landscapes, hands, and product shots simultaneously, and waiting for one to be is a way of not shipping. Ignoring audio until the end is the fourth, because it forces awkward visual compromises when the voice track needs a different rhythm than your edit.
No naming convention is the fifth mistake. Use a project underscore shot underscore take pattern from day one, and you will never again sort through files called final version two. Over-stylising before continuity is solved is the sixth. Nail the plain version first, then add grain, halation, and colour. Style applied to an inconsistent sequence just makes the inconsistency louder.
Planning Time and Output Realistically
Rough planning ratios for a one-minute narrative piece: brief and shot list one to two hours, reference collection one hour, generation three to six hours including retries, editing and sound three to four hours, quality control one hour. Expect roughly one accepted clip per three to six generations for complex motion, and closer to one in two for simple static shots. Track your own ratio for a week. It is the single best predictor of whether a project is feasible on your timeline.
Collaboration changes the maths. If two people work the same pipeline, split generation and post-production rather than splitting shots between them. Two people generating the same character with different references will produce two different people.
Budget contingency rather than optimising for best case. A project planned at exactly the optimistic estimate will slip the first time a model behaves unpredictably, which is often. Reserve twenty percent of your schedule for reshoots on hero moments.
FAQ
How many shots do I need for a 60-second video?
Between 12 and 20 for a paced narrative, or 8 to 12 if you favour long, meditative takes. Under 8 shots the edit feels static; over 25 and every shot becomes a flash.
Do I need editing experience to work this way?
You need basic timeline skills: trim, ripple delete, audio levelling, and colour balance. None of that is exotic, and an afternoon with any mainstream editor covers it. The AI-specific skills are prompt structure and reference discipline.
How do I stop characters from changing between shots?
Use one clean frontal reference per character, keep wardrobe descriptions identical across prompts, and reuse accepted frames as first frames rather than regenerating from text.
Is image-to-video always better than text-to-video?
No. It is better when identity or product fidelity matters. Text-to-video is faster and often more dynamic for atmosphere, transitions, and anything without a recurring subject.
What resolution should I generate at?
Generate at the highest resolution your workflow can afford in time, then upscale the accepted takes. Downscaling a good clip is trivial; recovering detail from a soft one is not.
How do I keep a project on schedule?
Cap retries per shot, keep a buffer shot list, and lock your edit structure before generating hero shots. When a shot exceeds its retry cap, use an alternative from the buffer list and move on.
Can I reuse one generation in several projects?
Yes, and you should. Tag accepted clips by location, mood, and subject so they are searchable. A library of vetted establishing shots saves hours on every future project.
Building a Durable Pipeline
Model names change every few months; the pipeline does not. Brief, references, shot list, model selection, prompt discipline, take logging, sound, grade, and quality control. Once those stages are habits, new models become upgrades you slot into a known process instead of experiments that reset your progress.
Start small. Produce one minute of finished video, measure how long it actually took stage by stage, and optimise the slowest stage first. Most creators discover that their bottleneck is not generation quality at all, but the absence of a shot list and a take log. Fix those two and the rest of the pipeline gets faster on its own.

