Start With the Deliverable, Not the Tool
Every AI video project begins with the same temptation: open a generator, type something evocative, and see what comes back. That approach can produce a striking clip, but it rarely produces a finished video. The gap between a demo and a deliverable is specification. Before you choose any model, write down the contract for the piece.
A practical deliverable spec answers seven questions:
- How long is the final video, and how many shots does that imply?
- What aspect ratio, resolution, and frame rate does the destination platform expect?
- Does it need voiceover, music, ambience, or dialogue?
- Are there on-screen text, logo, or lower-third requirements?
- What tone references exist that everyone can actually look at?
- How many review rounds are available, and who signs off?
- What is the hard deadline, and how much generation time does that leave?
Here is a concrete example. A 45-second teaser for a fitness app: vertical 9:16, 1080x1920, 24 fps, roughly 14 shots averaging 3.2 seconds, no on-screen text except a final logo lockup, one female voiceover that sounds energetic but not shouty, a music bed between 110 and 120 BPM, and two revision rounds. That single paragraph removes more wasted generation than any prompt trick you will learn later.
The wrong first question is which model to use. The right first question is what each shot must accomplish. Model choice is downstream of that decision, not upstream.
Why constraints accelerate generation
Unlimited options are expensive in time. Once aspect ratio, shot length, and audio plan are fixed, the plausible approaches narrow to a handful. You stop evaluating tools in the abstract and start evaluating whether a given approach satisfies a known requirement. Teams that specify first typically reach a locked cut in fewer iterations, because every rejection has a reason attached to it rather than a vague sense that something feels off.
Designing the Pipeline: From Script to Shot List
A shot list is the bridge between a written idea and a generated frame. It converts prose into discrete, generateable units, each with its own purpose and continuity requirements.
Beat sheet and shot intent
Start with a beat sheet: five to nine beats that carry the story. Then expand each beat into shots. For every shot, record six fields:
- Intent: what the shot must communicate to the viewer
- Shot size and camera move: wide, medium, close, push in, orbit, handheld
- Duration: how long it plays on screen
- Subject action: one clear, describable motion
- Lighting and mood: time of day, key direction, color temperature
- Continuity anchors: wardrobe, props, location details that must survive into neighbouring shots
A shot with two competing actions almost always generates worse than two shots with one action each. If a field is hard to fill in, the shot is probably underspecified rather than ambitious.
| Shot | Intent | Size and move | Duration | Continuity anchor |
|---|---|---|---|---|
| 1 | Establish city at dawn | Wide, slow push | 4.0s | Blue-grey palette, wet streets |
| 2 | Introduce runner | Medium, tracking | 3.0s | Teal jacket, same street lamps |
| 3 | Effort and focus | Close on face, handheld | 2.0s | Sweat, breath fog, same light direction |
| 4 | Product reveal | Insert, locked off | 1.5s | Logo lockup, clean background |
Asset and reference planning
Before generating, collect the raw material that will keep the sequence coherent: character reference images, wardrobe and prop references, location plates, and a small set of style frames. Name them consistently, for example character-runner-front, character-runner-side, style-dawn-city-01. A naming convention feels bureaucratic until the first time you have to regenerate shot nine and cannot remember which reference produced the version the client approved.
Matching Models to Shot Types
Different generation approaches solve different problems. Treating them as interchangeable is the most common source of rework.
Text-to-video, image-to-video, and video-to-video
Text-to-video is best for exploration, establishing shots, abstract transitions, and anything where exact framing is negotiable. It is fast and flexible but the weakest at holding a specific character or layout across takes.
Image-to-video starts from a still you control. Because composition, wardrobe, and colour are already fixed, it is the workhorse for character-driven sequences, product shots, and any frame the client has approved as a key visual. The trade-off is that you must first produce or source the still, which adds a step and a dependency.
Video-to-video and motion-transfer approaches take existing footage or animation and restyle it. They are ideal when the motion performance matters more than the surface look: dance, sport, gesture-driven scenes, or turning previz into a finished aesthetic. They require source footage with clean motion, which is a real production cost.
When to add upscaling and frame interpolation
Generate at the resolution and frame rate that suits the model, then finish separately. Upscaling helps when the final delivery is 4K or when a shot will be pushed in during the edit. Frame interpolation helps when a shot needs smoother slow motion, but it can introduce warping on fast limbs and complex textures. Always review interpolated shots at full speed and at half speed, because artefacts often hide in motion and appear when paused.
| Shot need | Recommended approach | What to verify |
|---|---|---|
| Establishing landscape | Text-to-video | Horizon stability, no drifting geometry |
| Consistent character | Image-to-video from reference | Face shape, wardrobe, hair length |
| Restyled live action | Video-to-video | Motion fidelity, edge tearing |
| Hero product beauty shot | Image-to-video, locked camera | Label legibility, reflections |
| Fast cutaway montage | Text-to-video, short clips | Palette match with neighbours |
Prompting for Motion and Camera Control
Prompts for video are not descriptions of a picture. They are instructions for change over time. The most reliable structure has five parts, written in this order.
The five-part prompt
- Subject: who or what, with two or three distinguishing details.
- Action: one primary motion, plus any secondary motion that must not conflict.
- Camera: framing, height, and movement, stated explicitly.
- Light and grade: direction, quality, time of day, colour bias.
- Constraints: duration, pacing, and anything that must not appear.
An example: a middle-aged runner in a teal windbreaker, action: steady forward stride with arms relaxed, camera: medium tracking shot at chest height moving with the subject, light: overcast dawn, cool blue-grey grade with soft rim light, constraints: no on-screen text, no other people, natural 24 fps cadence.
Notice how much of the prompt is camera and light, not subject. Motion language determines whether a clip is usable in an edit far more than the description of the character does.
Failure modes and fixes
- Morphing limbs and extra fingers: reduce the number of simultaneous actions, move the camera closer, and generate at a slightly larger framing then reframe in the edit.
- Unwanted cuts or scene changes: shorten the target duration and state a single continuous take explicitly.
- Jitter and shimmer on fine textures: avoid repetitive patterns such as fences, crowds, and foliage at distance, or use a shallower depth of field.
- Text artefacts: never rely on generated text. Add typography in post, where you control kerning, language, and legal wording.
- Default slow motion: many models drift toward a dreamy cadence. Specify natural speed and, if available, a shutter or motion amount setting.
Keep a running log of prompts that worked, with the model, settings, and seed where available. Your prompt library becomes the most valuable asset in the pipeline, far more than any single generation.
Holding Consistency Across an Entire Sequence
Consistency is what separates a sequence from a collection of clips. It has three layers: character, style, and scene.
Character reference packs
Build a pack of four to six images per character: front, three-quarter, profile, and one full-body with the exact wardrobe used in the story. Keep lighting consistent between reference images, because inconsistent references teach the model inconsistent appearance. When a character appears in multiple shots, generate those shots in the same session with the same references rather than returning to them days later.
Style bibles and colour scripts
A short style bible prevents drift. Write down the palette, contrast level, film grain or digital cleanliness, lens character, and what the piece should never look like. A colour script assigns a dominant colour to each beat: cool blues for the setup, warmer amber for the turn, and clean neutrals for the product moment. When two adjacent shots clash, the colour script usually tells you which one is wrong.
Scene continuity and shot-to-shot matching
Check three things between neighbouring shots: light direction, palette, and motion energy. If shot three moves left and shot four moves right, the cut can still work, but the reversal should be intentional. Continuity errors that audiences notice are rarely about props; they are about light and speed.
Assembly: Editing, Sound, and Finishing
Generation produces material. Editing produces a video. Budget as much time for assembly as for generation, if not more.
Cut rhythm and pacing
AI clips often look best when they are cut shorter than the generated duration. Trim to the moment of clarity: the frame where the action reads instantly. A common mistake is letting every shot run long enough to expose its weaknesses. Cut on motion, cut on a beat in the music, and vary shot length so the sequence breathes instead of humming along at one tempo.
Voice, music, and ambience
Voiceover locks the emotional register. Generate or record it early, before the final edit, so you cut to the performance rather than retrofitting audio to picture. Music provides the tempo grid; ambience hides the small sonic gaps that make generated footage feel synthetic. Layering three elements, voice, music, and a light ambience bed, is usually enough for a short piece. Keep dialogue sparse: lip-synced generated dialogue remains the least reliable part of the stack, and a voiceover over a face that is not speaking reads as intentional, while a badly synced mouth reads as broken.
Grade, aspect ratios, and delivery specs
Grade the sequence as a whole, not shot by shot. A single adjustment layer for contrast and colour balance goes a long way toward making disparate clips feel like one film. Then export each required aspect ratio separately, checking safe areas for text and logos. If you deliver vertical and horizontal versions, reframe deliberately rather than relying on automatic cropping.
Quality Control: Checklists and Review Loops
A structured review catches the mistakes that a tired eye misses. Run every shot through the same short checklist before it enters the timeline.
- Does the action read in one viewing without explanation?
- Is the face stable throughout, including at the end of the clip?
- Do hands, feet, and any held objects behave?
- Is the light direction consistent with neighbours?
- Does the grade match the colour script?
- Are there text artefacts, watermarks, or unwanted logos?
- Is the motion speed natural for the intent?
- Does the shot survive a full-screen viewing, not just a thumbnail?
Build in a review gate after rough assembly and before sound design. Fixing a shot at the rough-cut stage costs minutes; replacing it after the client has heard the music costs a day.
Troubleshooting common problems
| Symptom | Likely cause | Practical fix |
|---|---|---|
| Shot feels synthetic | Uniform lighting, no grain, over-clean motion | Add texture references, slight grain, varied camera energy |
| Character changes mid-sequence | Inconsistent references or separate sessions | Rebuild the reference pack, regenerate in one session |
| Sequence feels disjointed | Mixed palettes and motion directions | Apply a unifying grade and reorder for motion continuity |
| Clip looks mushy at 4K | Generated at lower resolution | Upscale and re-sharpen selectively, check fine detail |
| Audio feels thin | Single element only | Layer voice, music, and ambience separately |
Scaling Production: Templates, Batches, and Handoffs
Once a workflow works once, the goal is to make it repeatable without making it rigid.
Reusable presets and naming conventions
Save prompt templates by shot type, not by project: establishing wide, character medium, product insert, transition. Store approved reference packs and grade recipes in one place. Adopt a naming convention that encodes project, shot number, and version, for example fitness-teaser-s03-v04. Versioning is the cheapest insurance in a collaborative pipeline.
Batching and review gates
Generate in batches grouped by shot type rather than by story order. Generating all close-ups together keeps the reference packs loaded and makes comparisons immediate. Insert a review gate after each batch: approve, revise, or reject, then move on. Unreviewed batches pile up into a mountain of clips nobody can evaluate clearly.
Working with collaborators
Handoffs fail when context lives in someone's head. A shared document with the shot list, reference packs, approved takes, and open questions keeps a second editor productive on day one. Write the document for a person who has never seen the project, and it will serve you just as well three weeks later.
Ethics, Rights, and Client Communication
AI video work carries obligations that traditional production does not. Address them early, in writing.
- Likeness and voice: never generate a recognisable person or voice without documented permission.
- Training and source material: know the terms of the tools you use, and ask before using client footage as a style reference.
- Disclosure: agree with the client whether and how AI involvement is disclosed to the audience.
- Expectations: explain that generation is probabilistic. Show two or three representative takes early so nobody is surprised by variability later.
- Scope: define how many revision rounds are included, because unlimited revision on generative material is unbounded work.
A short written summary of these points avoids most disputes and often becomes a selling point, since it signals that you have done this before.
FAQ
How many shots can I realistically finish in a day?
For a short branded piece, plan on six to ten finished seconds per hour of focused work once your references and prompts are ready. Exploration days produce less. Consistency, not speed of generation, is usually the limiting factor.
Should I generate at final resolution or upscale afterwards?
Generate at the resolution where the model gives stable results, then upscale for delivery. Upscaling is cheaper than repeatedly gambling on high-resolution generations that may drift in composition.
Why do my characters change between shots?
Almost always because references changed, or because shots were generated in separate sessions with different settings. Rebuild a four-to-six image reference pack, then regenerate the affected shots together in one sitting.
Is image-to-video always better than text-to-video?
No. It is better when composition and identity must be controlled. Text-to-video remains faster for exploration and for shots where exact framing does not matter, such as establishing landscapes and abstract transitions.
How do I make AI footage feel less artificial?
Three levers: add texture and imperfection, vary camera energy between shots, and cut shorter than you think you need to. Uniform lighting and perfectly smooth motion are the strongest tells.
Can I use generated dialogue?
For short lines in tight framing it can work, but it is the least reliable element. Prefer voiceover, off-screen dialogue, or shots where the speaker is not clearly visible. Always review lip sync at full speed before committing.
What should I do when a client rejects a batch?
Ask which specific field failed: action, framing, light, or identity. A vague rejection cannot be fixed with a new prompt. Once you know which field is wrong, change only that field and regenerate, keeping everything else identical so you can compare properly.
Do I need a shot list if I am working alone?
Especially then. Working solo removes the meeting where assumptions get caught. A one-page shot list takes twenty minutes and routinely saves a full day of regeneration.
How do I keep a long project from drifting stylistically?
Freeze a style bible and colour script after the first rough assembly, then hold every new shot against them. When a shot disagrees with the bible, change the shot, not the bible, unless the client has approved a deliberate shift.
What is the most common workflow mistake?
Generating before specifying. Nearly every downstream problem, from inconsistent characters to unusable aspect ratios, traces back to starting the first clip before the deliverable was written down.
A workflow is not a set of tools; it is a sequence of decisions with checkpoints. Specify the deliverable, plan the shots, choose approaches per shot type, prompt for motion, protect consistency, assemble with sound, and review against a checklist. Do that consistently and the technology stops being a gamble and starts being a production method.


