How text and image prompts divide the work
Generative video has two entry points, and the choice between them shapes everything downstream. Text-to-video starts with language: you describe a scene, and the model invents framing, blocking, lighting, motion, and often the emotional tone. Image-to-video starts with a still you already control: you supply a frame, and the model animates it while trying to preserve composition and identity.
That difference matters more than any single generation feature. Text prompts give you speed and surprise. You can sketch a sequence of ten shots in an afternoon without drawing anything. But because the model makes compositional decisions for you, matching one shot to the next becomes genuinely hard. Image prompts give you continuity and art direction. Your character looks the same, your palette stays locked, your product packaging does not mutate between cuts. The tradeoff is preparation time and less freedom in how motion unfolds.
In practice, most finished work uses both. You build or generate keyframes as stills, approve them the way you would approve storyboard panels, then animate them. Text-to-video fills gaps, generates B-roll, and explores alternate ideas cheaply. A hybrid pipeline looks like this: write the beat, generate or photograph the keyframe, animate the keyframe, generate two or three text-only alternates for the same beat, then choose the best take in the edit.
The teams that get consistent results treat generation as one stage in a production pipeline rather than a magic button. They storyboard first, they write prompts with the discipline of a shot list, and they review output against explicit criteria instead of vibes.
Choosing your generation mode
Before writing a single prompt, decide which of the four modes each shot needs. Mixing modes inside one project is normal, but you want the decision to be deliberate.
Text-to-video
Use it for establishing shots, abstract transitions, environments, crowd scenes, weather, and anything where the exact composition is negotiable. It is also the fastest way to test whether an idea works at all. Cost of failure is low because you have invested nothing but a paragraph.
Image-to-video
Use it when identity, product detail, typography, or layout must survive the shot. Character close-ups, product hero shots, branded environments, and any frame that has to match a previous frame belong here. Feed the model a clean, well-lit still at the target aspect ratio. Soft or noisy source images produce soft or noisy motion.
Hybrid keyframe animation
Generate a first frame and a last frame as stills, then let the model interpolate between them. This is the closest thing generative video has to traditional animation blocking, and it is the most reliable way to hit a specific action on a specific beat. It works especially well for reveals, transformations, and camera moves with a defined end position.
Video-to-video and restyle
When you already have live-action footage, restyling or extending it is often more efficient than generating from scratch. Use it for stylistic passes, background replacement, cleanup, frame-rate smoothing, and extending a clip past what the camera captured. Keep an unmodified master of every source clip.
A quick decision rule: if the shot must match something, start from an image or existing footage. If the shot must evoke something, start from text.
Plan the shot list before you open a tool
Generative tools reward preparation disproportionately. A twenty-minute planning pass saves hours of rerolling.
Start with a beat sheet: five to twelve story beats, each one sentence. Then convert beats into a shot list with one row per shot and these columns:
- Shot ID and beat reference
- Mode (text, image, hybrid, restyle)
- Duration target in seconds
- Subject and action in plain language
- Camera: framing, height, movement, lens feel
- Lighting and time of day
- Palette and grade notes
- Audio: dialogue, ambience, effect, music cue
- Continuity anchors: wardrobe, props, location, character reference
- Status and approved take
Two planning decisions carry most of the value. First, keep generated shots short. Three to six seconds per shot is a comfortable working range; longer generations drift, morph, and lose coherence. You can always extend in the edit by cutting between takes. Second, define the camera move in words before generating, because models interpret vague camera language inconsistently. "Slow push in from medium to close" is usable. "Dynamic camera" is not.
Finally, decide your delivery spec now: aspect ratio, resolution, frame rate, and total runtime. Generating vertical 9:16 when the client needs 16:9 wastes every take you produced.
Prompt architecture that survives model swaps
Different generators weight prompt terms differently, but a consistent internal structure lets you port one prompt across tools with minimal rewriting. Build every prompt from five layers.
Subject, action, and environment
Lead with the concrete. Who or what, doing what, where, and in what condition. Avoid adjectives that describe your opinion rather than the frame.
A ceramicist in a grey apron shaping a bowl on a wooden workbench,
studio window behind her, clay dust in the air, morning light
Camera and lens language
Specify framing, height, movement, and pace. Mention the lens feel only if it changes the image meaningfully.
Medium shot at chest height, slow dolly right, shallow depth of field,
subject stays centre frame
Lighting and grade
Describe direction, quality, and color temperature. This single layer fixes more inconsistencies than any other.
Soft north-facing window light, cool shadows, warm highlights,
muted desaturated palette with a slight teal cast in the shadows
Motion and physics
Say what moves and how fast, and what should stay still. Models love to add motion when told nothing, so explicit stillness instructions are useful.
Fabric drifts gently, wheel rotates slowly, background remains static,
no camera shake, no subject morphing
Style anchoring and exclusions
Anchor style with reference terms or, better, with your approved keyframe. Then add short negative instructions for known failure modes: no text overlays, no extra fingers, no sudden cuts, no lens flare unless requested.
Keep prompts under roughly 120 words. Long prompts dilute attention across too many concepts, and later clauses get ignored. If you need more control, generate in stages and edit rather than overloading one prompt.
Model selection criteria without brand loyalty
The market changes monthly, so choose models per shot rather than per project. Score candidates against these criteria:
| Criterion | What to test | Why it matters |
|---|---|---|
| Motion realism | Walking, hands, fabric, water, crowds | Determines whether you can use the shot at all |
| Duration ceiling | Coherence at 4s, 8s, 12s | Longer windows reduce the number of takes |
| Resolution and frame rate | Native output before upscaling | Affects delivery and crop headroom |
| Controllability | Camera control, keyframes, masking, seeds | Enable fixes without full regeneration |
| Identity retention | Same face across three shots | Decides if you need a reference-image pipeline |
| Audio support | Dialogue, ambience, lip sync | Can remove a separate production stage |
| Latency and queue time | Time to first usable take | Sets your realistic iteration count |
| Commercial licensing | Terms for your use case and territory | Protects client delivery |
| Export and metadata | Codec, alpha, color handling | Prevents downstream conversion problems |
Run a one-shot bake-off before a project starts. Take your hardest shot, generate it in three or four tools with functionally identical prompts, and compare the results side by side in a timeline. Keep the winner for hero shots and the cheapest acceptable option for filler.
Model churn is a feature, not a problem, if your pipeline keeps prompts, keyframes, and edit decisions in formats that outlive any single tool. Store prompts as text files in the project folder. Store approved keyframes with version numbers. Never let a proprietary project file become the only place a decision exists.
End-to-end workflow from brief to approved sequence
This is the loop most teams converge on after their first few productions.
- Lock the brief. Runtime, ratio, tone, audience, and the single idea the piece must land.
- Write the beat sheet and shot list. Approve them before generating anything.
- Build a look board. Six to ten reference images that define palette, lighting, and framing.
- Create keyframes. Generate or photograph one approved still per narrative shot. Reject weak keyframes immediately; animating a mediocre frame rarely improves it.
- Write prompts per shot. Use the five-layer structure and keep a shared vocabulary file so the whole team describes the same look the same way.
- Generate in batches of two or three takes. Compare, then adjust one variable at a time. Changing five things between attempts teaches you nothing.
- Select and log. Name approved files with shot ID and take number. Note the seed and prompt version.
- Assemble a rough cut immediately. Do not polish takes in isolation. Rhythm problems hide inside individual clips.
- Fix in the edit first. Speed ramps, reframes, split screens, and cutaways solve more defects than regeneration.
- Add audio, grade, and delivery masters. Export a textless master, a captioned version, and platform-specific cuts.
The most common process failure is skipping step eight. Teams generate forty beautiful clips and discover in the assembly that the sequence has no pace. Rough cut early, regenerate late.
Keeping characters and style consistent across shots
Consistency is the hardest problem in generated video, and it is solved with references, not adjectives.
Build a character sheet. One front-facing still, one three-quarter, one profile, at neutral expression and consistent lighting. Animate from these rather than from new text descriptions each time. If your tool supports reusable subject references, register the character once and reuse it.
Reuse the first frame of the previous shot. Starting a new shot from the last frame of the old one is the cheapest continuity trick available, and it works across tools.
Lock the palette numerically. Decide on three or four values for shadows, midtones, highlights, and accent, then apply a consistent grade in the editor. Small grade differences read as separate scenes even when the content matches.
Fix wardrobe and props in language. Describe garments with material, color, and cut rather than a single noun. "Charcoal wool overcoat with a wide collar" survives regeneration far better than "dark coat."
Keep shot grammar uniform. If the sequence uses chest-height medium shots with slow lateral moves, do not insert one handheld low-angle unless it is a deliberate accent. Consistency of camera language reads as authorship.
Accept managed imperfection. Photorealistic faces across many shots will drift slightly. You can hide a great deal with framing, motion blur, shadows, and cuts on action. Choose the shot where drift is least visible rather than chasing a perfect match indefinitely.
Audio, pacing, and the edit
Generated picture is only half the deliverable. Audio is where most AI video feels amateur, because creators treat it as an afterthought.
Work sound-first for anything with dialogue or narration. Generate or record the voice track, set the timing, then generate shots to fit the audio. Shot durations that feel odd in isolation often read perfectly against a voice track, and the reverse is equally true.
For each shot, decide what carries the scene: dialogue, ambience, or music. Do not let all three compete. Layer ambience underneath at low level for spatial reality, place effects on the action beats, and duck music under speech. If the tool generates audio with picture, treat it as scratch and replace it during the mix.
Pacing rules that hold up in generative work:
- Cut on motion. A subject reaching, turning, or stepping hides soft endings.
- Vary shot length deliberately. Uniform three-second cuts feel mechanical by the fourth shot.
- Give every act a longer hold. One slow shot per sequence resets attention.
- Use transitions sparingly. Hard cuts age better than stylized wipes.
- End on a frame you would print. The last shot should survive a pause.
For extensions and speed changes, do them in the editor rather than asking for a longer generation. A 4-second clip stretched to 6 seconds with a light optical flow pass, cut with two inserts, reads as a 10-second sequence and costs nothing extra to produce.
Troubleshooting common defects and QC checklist
Most failures are predictable, and most have an editorial fix that costs less than regeneration.
Face warping mid-shot. Reduce shot length, lower motion intensity, and start from a higher-resolution still. If the warp lands at the end, cut earlier and use the clean portion.
Hand and finger artifacts. Reframe so hands leave the frame, place them behind props, or switch to a wider shot where the detail is not readable. Avoid close-up hand action unless the model handles it reliably.
Temporal flicker and texture crawl. Often a resolution or compression issue. Regenerate at higher native resolution, then use a light temporal denoise. Beware over-denoising, which produces a plastic look.
Camera drift and unwanted zooms. Write explicit stillness and framing constraints. If drift persists in image-to-video, use a keyframe pair that holds the composition at both ends.
Character identity shift between shots. Return to the character sheet and animate from reference stills. Add a grade pass to unify tone, and avoid shots where the face is both large and in motion.
Morphing objects and melting geometry. Shorten the shot, simplify the action to one verb, and remove secondary motion from the prompt. Complex simultaneous motion is where coherence collapses first.
Text and logo artifacts. Never generate text you need to read. Composite real typography in the editor over a clean plate.
Lip sync drift. Generate shorter dialogue segments, keep the face at a consistent angle, and align in the editor on plosive sounds rather than word starts.
Before you call a sequence finished, run this check:
- Does every shot have a clear subject and a single action?
- Do faces, wardrobe, props, and locations hold across cuts?
- Is the lighting direction consistent within a scene?
- Are there any readable artifacts at full resolution, on a large screen?
- Does the audio mix survive a phone speaker and headphones?
- Are captions accurate and timed to speech?
- Do exports match the delivery specs, including color space and audio loudness target?
- Is every licensed asset documented in the project folder?
FAQ
How long should a generated shot be?
Three to six seconds works for most narrative content. Shorter clips stay coherent and give you more editorial control. Reserve longer generations for slow, low-motion shots where drift is less visible.
Is text-to-video or image-to-video better for product work?
Image-to-video, almost always. Products need shape, label, and color accuracy. Start from a clean studio still at the delivery aspect ratio and keep motion minimal.
How many takes should I expect before a usable shot?
Budget two to four attempts for straightforward shots and six or more for anything involving faces in motion, hands, or complex physics. Track which prompt layers you changed so the process compounds instead of repeating.
Can I mix output from several tools in one project?
Yes, and most teams do. Normalize resolution and frame rate on import, apply a single grade at the sequence level, and keep a note of which tool produced each shot for future reference.
What is the fastest way to fix a shot that almost works?
Fix it in the edit first. Trim the strong portion, reframe, add a cutaway, or slow it down. Regenerate only when the defect spans the whole clip.
Do I need a storyboard for a thirty-second piece?
Yes. Even a rough six-panel board prevents the most expensive mistake in generative video: producing attractive clips that do not assemble into a sequence.
How do I keep a series visually unified over many episodes?
Maintain a shared project kit: character sheets, palette values, lighting recipes, prompt vocabulary, grade settings, and audio beds. Reuse the kit rather than re-describing the look in every new session.
The core discipline is simple. Plan the sequence, control the frame before you animate it, generate short takes, cut early, and treat models as interchangeable tools inside a workflow you own.


