Start With the Deliverable, Not the Model
Almost every disappointing AI video project fails before the first frame is generated. The model was fine, the prompts were plausible, and the hardware was adequate. What was missing was a one-page brief that stated what the finished piece had to accomplish and how anyone would know it worked.
Write that page first, and keep it to one page. It should answer what the deliverable is — a fifteen-second vertical teaser, a ninety-second explainer, a looping background clip for a landing page — plus the aspect ratios required, where the piece will be watched, the tone, the deadline, and the review chain. Add a short list of non-negotiables: the product mark must be legible, the presenter's likeness must match a supplied reference, the palette must follow brand guidelines, and no text should be baked into frames because captions get added later.
The brief matters more with generative tools than it did with cameras, because generation offers infinite options and imposes no constraints. A camera crew forces decisions: the location is rented for four hours, so you shoot. A prompt box forces nothing. Every unmade decision becomes an extra round of generation, and extra rounds are where schedules quietly die.
Two numbers in the brief do most of the work. Duration determines shot count, and shot count determines how much consistency labour you owe. A fifteen-second piece is three to five shots; a ninety-second piece is fifteen to twenty-five. Aspect ratio determines composition. A frame that sings in 16:9 often collapses in 9:16 when the subject sits in the middle third and everything else is cropped away. Decide crops before generating, and when multiple formats are needed, compose with generous headroom and keep key action away from the edges.
Finish with references. Collect eight to twelve images or clips and annotate each one specifically: this palette, this lens compression, this grain, this lighting direction. A vague vibe is not a reference; it is a wish. An annotated set turns taste into instructions, and instructions are what a model can follow.
Choosing a Model Stack: Decision Criteria
Most tool confusion comes from comparing products when the useful comparison is between categories. Four categories cover almost everything a small studio needs.
- Exploration models: fast, tolerant of vague prompts, weak on fine detail. Use them to find a direction in twenty minutes, never for final frames.
- Hero still models: high fidelity, strong control over texture and lighting. Use them for key art and for any frame you plan to animate.
- Motion models: tuned for temporal coherence, believable camera moves, and faces that hold their shape. Use them for short, controlled movement.
- Finishing tools: upscalers, interpolators, denoisers, colour and grain tools. These decide whether a technically adequate frame reads as professional.
The pattern that works best for small teams is approve the still, then move it. Generate a hero frame with a still model, get sign-off, then animate only the approved frame with a motion model. Randomness drops sharply because you are no longer gambling on composition and motion at the same time.
Evaluate new models with a fixed test set instead of a gallery of cherry-picked samples. Five prompts are enough: a portrait with difficult hair lighting, a wide landscape with layered depth, an interior containing legible signage, a fast action moment, and a low-light scene with practical lights. Run every candidate on the same five prompts at the same resolution and judge side by side. Tools that win on portraits frequently lose on text and architecture, and you will only discover that by testing your own material.
Three criteria matter more than benchmark scores. Reproducibility: can you fix a seed and get the same result tomorrow? If not, every approved frame is a one-off and continuity becomes guesswork. Iteration cost: how long does one variation take and what does it consume in money or machine time? Cheap iteration changes how boldly you experiment. Control surface: does the tool expose seed, guidance strength, reference images, negative prompts, motion amount, and camera controls? More levers mean less reliance on luck.
Record the answers in a half-page stack document: which model for what, which seeds are locked, which settings are house defaults. That page prevents the team from re-deciding the same questions on every project, and it makes onboarding a new collaborator a matter of reading rather than interviewing.
Prompt Architecture That Scales
Prompts are the source code of an AI video project. Treat them accordingly: fixed order, versioned, stored, reused.
The five blocks
Write every prompt in the same order so it can be debugged block by block.
- Subject — who or what, apparent age range, wardrobe, expression, one distinguishing detail.
- Action — the verb of the frame. One clear action beats three competing ones.
- Environment — location, era, weather, background density, what fills the frame edges.
- Camera — lens, angle, distance, movement, framing.
- Light and grade — source, direction, contrast, palette, grain or film emulation.
When a result fails, you can point at the block responsible and change one variable at a time. That is what turns iteration into engineering rather than superstition. It also means a junior collaborator can improve a prompt without understanding the whole pipeline.
Keep negative prompts short
Maintain a brief negative list for recurring defects: extra fingers, duplicated limbs, watermark text, melted background lettering, plastic skin. A short list removes obvious failures. A long list — thirty terms, half of them abstract — pushes the sampler away from so many directions that fresh artifacts appear. When a defect keeps returning, fix it in the positive prompt or in post production instead of stacking more negatives.
Seed discipline
Record the seed for every approved frame. Seeds are the cheapest continuity tool available: the same seed with a small prompt change behaves like a controlled variation on a known-good look. Lose the seed and you lose the ability to make a matching shot next week. This single habit separates teams who can extend a project from teams who start over.
Versioning and logs
Keep prompts in a plain text file or spreadsheet with columns for shot ID, model, version, seed, negative list, notes, and the filename of the approved render. Add a column for what changed and why. Six weeks later, when someone asks for one more shot in the same style, that file is the difference between an afternoon and a week.
Keeping Characters and Scenes Consistent
Consistency is the hardest part of AI video and the part most often skipped. It comes from anchors, not luck.
Identity anchors
Start with a reference image of the character in neutral light on a plain background. Feed it into every shot. Where the tool supports adapters or lightweight style training, train on ten to twenty images covering varied angles, distances, and expressions. Write wardrobe decisions down, because models happily reinvent a jacket between shots.
Name the anchors. A reusable block like lead — charcoal coat, red scarf, short hair can be pasted into every prompt, and it is far more reliable than re-describing the character from memory. Do the same for locations: kitchen — morning light from left window, pale oak surfaces, blue kettle on the second shelf. These named blocks become a private vocabulary that both the prompt author and the reviewer understand instantly.
A continuity checklist
- Same seed family for each location
- Same focal length for shots inside one scene
- Consistent time of day and light direction
- Wardrobe, props, and hair state tracked scene by scene
- Grade applied after assembly, not baked into every frame
- Screen direction and eyelines respected across cuts
- Background elements that recur, such as signs, furniture, and weather, logged in the shot file
Run the checklist before rendering a batch, not after. Re-rendering twenty clips because the light flipped from morning to dusk is expensive, and it is entirely avoidable.
The End-to-End Workflow, Stage by Stage
Stage one: brief and moodboard
Spend an hour writing the one-page brief and assembling annotated references. This hour prevents most revision rounds later. The output of this stage is a document, not a render.
Stage two: look development
Generate twenty to forty cheap, low-resolution stills. Do not chase polish. Approve two or three directions, then refine one. Lock palette, lens feel, and texture before generating anything that moves, because everything downstream inherits these decisions. If the look is wrong, nothing later can rescue it.
Stage three: shot generation
Work shot by shot, producing three to five variations per shot and selecting the strongest. Generate a few frames beyond each cut point so the editor has handles. Name files with scene, shot, and take numbers so the timeline assembles itself logically. Keep a reject folder; rejected frames often become the right frame for a later scene.
Stage four: motion, sound, and finishing
Animate approved frames, then assemble in an editor. Add ambience and music early, because sound changes pacing decisions more than most creators expect. Stabilise, grade, and export a review cut. Keep a three-line log of what changed between versions so nobody relitigates a settled decision.
Stage five: review and delivery
Watch the cut three ways: once with sound, once muted, once at double speed. Silent viewing exposes composition problems; fast viewing exposes pacing problems. Then deliver in every aspect ratio the channels need, with captions, thumbnails, and an archived note describing the workflow. That note becomes the starting point for the next project.
A Worked Example: Twenty Seconds, Four Shots
Setting: a twenty-second teaser for a coffee brand, 16:9 and 9:16, warm morning interior, one character. Generation volume: about twelve stills to lock the look, then three variations per shot, roughly twenty-four renders total plus a handful of retries.
Shot one, establishing wide: wide shot, sunlit kitchen, woman in a linen shirt pouring water into a glass carafe, 24mm lens, slow push in, practical window light from the left, warm neutral palette, soft contrast, fine grain. Negatives kept short: extra limbs, text, watermark, blown highlights.
Shot two, medium: same seed family, same light direction, distance changed only. Medium shot, chest up, same character at the same counter, pouring, 50mm lens, static camera, steam visible in backlight. Because seed and light direction stay identical, the shirt, hair, and counter match without retouching.
Shot three, close-up on hands: close-up, hands and carafe, water catching the light, 85mm lens, slight handheld drift, same lighting. Hands are the classic failure point, so generate several variations and choose the one where fingers read correctly at thumbnail size. If none of them do, cover the shot with a wider framing rather than fighting the model.
Shot four, product beat: the cup on the counter, morning light across the rim, slow tilt up to reveal the window, 50mm, same grade. End on a frame with clean negative space for the logo, composed for both aspect ratios.
Post: assemble the four shots, add room tone plus a low warm pad, grade once across the whole sequence, stabilise the handheld drift, and export both formats. Note how the lens progression from 24mm to 50mm to 85mm reads as a natural shot sequence. Consistent framing logic hides small inconsistencies in detail, which is why lens discipline is a continuity technique and not just a stylistic preference.
Planning Time, Hardware, and Roles
Planning numbers beat opinions. For a one-minute narrative piece with roughly fifteen shots, allow two to three days for look development, one day per three to four shots for generation and selection, and one to two days for motion, sound, and finishing. Add buffer for one round of notes, and assume that round touches about a third of the shots.
On hardware, a modern consumer graphics card with plenty of video memory handles local still generation comfortably. Video generation is heavier, so batch overnight and keep the machine free during the day for editing and review. If nobody on the team enjoys driver maintenance, run exploration on hosted tools and reserve local generation for shots that need exact seeds or private material.
Roles matter more than tools at this scale. Even a two-person team should separate the person who writes prompts from the person who reviews them. The prompt author optimises for the model; the reviewer optimises for the audience. When one person does both, the review step quietly disappears and the final cut suffers for it.
Budget in three currencies rather than one: machine time, human attention, and revision rounds. Attention is the scarce resource. Every extra variation you generate costs review time, so cap variations per shot and spend the savings on a second look at the edit, where problems are cheaper to fix.
Troubleshooting: Common Failure Modes and Fixes
- Faces drift between shots. Cause: no identity anchor and no locked seed. Fix: reintroduce the reference image and rebuild the scene from one seed family.
- Hands and teeth look wrong. Cause: resolution too low for fine structure, or motion too ambitious. Fix: generate the still larger, repair hands in post if necessary, and reduce motion amount.
- The scene looks generic. Cause: generic input. Fix: define a palette, a lens set, and a lighting recipe, then reuse them instead of randomising every prompt.
- Text in the frame is unreadable. Cause: generative models still struggle with typography. Fix: leave clean space and add type in the editor.
- Motion looks like a slideshow. Cause: too few frames of genuine movement, or an interpolator doing all the work. Fix: keep moves small, raise motion consistency, and cut between short clips rather than stretching one long one.
- Colour shifts between shots. Cause: grading baked into individual generations. Fix: generate as neutrally as possible and grade once across the assembled sequence.
- Everything takes too long. Cause: no fixed brief, so the target keeps moving. Fix: freeze the brief, approve the look, then generate to a shot list.
- The same prompt gives a different result every time. Cause: unfixed seed, or a service that updates its models silently. Fix: pin versions where possible and log the model version beside every approved frame.
Rights, Consent, and Review
Wider creative latitude raises the bar on your own process rather than lowering it. Keep a checklist and keep the paperwork with the project files.
- Written permission for any real person's likeness, especially in commercial work
- No trademarked logos or protected characters in paid output
- No sexualised depiction of minors in any form, ever
- No defamatory or misleading portrayals of identifiable people
- Clear disclosure where platform rules or local law require it
- A second reviewer for anything touching politics, real events, or sensitive subjects
- A record of which model version produced which asset, in case a licence question arrives later
Documentation protects a studio more than any filter does. Store prompts, references, permissions, and model versions somewhere that outlives the current laptop, and treat that archive as part of the deliverable rather than an afterthought.
FAQ
How many variations should I generate per shot?
Three to five. More rarely improves selection and always increases review time.
Do I need an expensive workstation?
Only for local generation. Hosted tools run on modest hardware, and the trade is that you accept someone else's policy layer and availability.
What is the biggest time saver in an AI video pipeline?
Prompt versioning combined with seed tracking. Together they turn guesswork into iteration, and iteration is what actually improves a film.
Should I train a custom character or style?
If a project has more than a dozen shots featuring the same character or look, yes. Below that, reference images and locked seeds are usually enough.
How do I keep a series consistent across episodes?
Freeze a style document: palette, lens set, grade recipe, and exact prompt blocks for recurring characters and locations. Every episode starts from that document instead of a blank page.
Can AI video replace a live shoot?
For some formats, yes: product inserts, abstract backgrounds, concept films, motion graphics. For dialogue-driven scenes with subtle performance, live action still wins, and hybrid pipelines — generated backgrounds with real footage composited over them — are often the pragmatic answer.
How do I judge whether output is good enough to ship?
Watch the assembled cut, not individual frames. A frame that looks weak in isolation often works fine in motion, and a frame that looks stunning often breaks the moment it moves.
What should I do when a hosted tool refuses a prompt I consider legitimate?
Test the same prompt on a different engine to confirm where the block sits, then decide: rewrite for the classifier, switch tools, or drop the shot. Spending an hour fighting an opaque filter is rarely the best use of that hour.
Why does my second episode look nothing like the first?
Usually because the prompts were rewritten from memory. Reuse the archived prompt blocks and the locked seeds, and change only what the new story requires.
A calm, documented pipeline beats a clever one. Choose tools by category, write the brief before the prompt, approve the still before the motion, track seeds and versions, review with sound and without, and archive everything. Do that consistently and the technology stops being a gamble and starts behaving like a craft.


