Why text and image prompts belong to the same workflow
Most creators discover generative video through one of two doors. Either they type a sentence into a text-to-video generator and hope the result resembles the picture in their head, or they upload a still image and ask a model to make it move. Both doors lead to the same room: a finished sequence in which every shot has to match the shot before it.
Treating text-to-video and image-to-video as competing techniques is the first mistake. They solve different problems:
- Text-to-video is for shots that do not exist yet. Establishing shots, abstract transitions, B-roll, dream sequences, anything where you have no visual reference and no obligation to match an existing frame.
- Image-to-video is for shots that must connect to something already approved. Product hero shots, character close-ups, storyboard panels, photos you shot yourself, brand assets that must not drift.
A healthy project uses both. You might generate a hero keyframe as an image, animate it with image-to-video to guarantee the composition, then cut to a text-to-video insert for a transition that carries no continuity burden. The pipeline below is built around that division of labour, and it works whether you are producing a 15-second social ad or a five-minute explainer.
The goal of this guide is not to sell you a specific platform. It is to give you a repeatable process: how to choose a model category for each shot, how to prompt motion rather than appearance, how to hold characters together across cuts, and how to finish the result so it does not look like a demo reel.
Decide these five things before you generate a single frame
Generating first and planning later is how budgets evaporate. Spend twenty minutes on decisions before you touch a prompt field.
Clip length and motion complexity
Short clips of two to five seconds hide more defects than long ones. A model that produces a beautiful three-second shot may fall apart at eight seconds as the subject's face drifts or limbs multiply. Plan your edit around short shots glued together with cuts, match cuts, and sound bridges rather than chasing single long takes.
Ask of each shot: does the camera move, does the subject move, or both? Both at once is the hardest case. If the subject must walk while the camera tracks, expect more iterations than a static frame with drifting light.
Reference fidelity versus motion freedom
Image-to-video forces the model to respect your starting frame, which is exactly what you want for brand-critical content. The trade-off is that aggressive motion can warp the reference. If you need a product label to stay legible, keep motion subtle: a slow push-in, a gentle parallax, drifting light. Save the dramatic camera work for text-to-video shots where nothing has to stay recognisable.
Resolution, aspect ratio, and delivery format
Decide the master aspect ratio early. Vertical for short-form feeds, 16:9 for YouTube and presentations, 1:1 or 4:5 for certain ad placements. Generating in the wrong ratio and cropping later throws away detail and often cuts heads. Where a model supports only one ratio, generate at the largest supported size and reframe in the edit rather than the other way around.
Audio strategy
Music, voice-over, and sound effects carry more of the perceived quality than most people expect. Decide now whether you will use a synthetic voice, record your own, or work without narration and let music and on-screen text do the work. Voice-over also dictates pacing: a 40-word line takes roughly 15 seconds, so your shot count follows from your script.
Continuity obligations
List every element that must not change: a character's jacket, a logo, a room's layout, the direction of sunlight. This list becomes your review checklist later. Projects rarely fail because a shot looks bad; they fail because shot four and shot five clearly come from different universes.
A six-stage production pipeline you can repeat
The workflow below scales from a solo creator to a small team. Each stage has a clear exit condition, which keeps you from polishing stage two while stage four is still unstarted.
Stage 1: Script to beat sheet
Write the script, then reduce it to a beat sheet: one line per shot describing what the viewer must see and feel. For a 60-second piece you will typically land between 12 and 20 beats. Mark each beat with one of three tags:
- Key — hero imagery, must match approved references. Route to image-to-video.
- Texture — B-roll, atmosphere, transitions. Route to text-to-video.
- Graphic — text, charts, UI, lower thirds. Route to a standard editing tool, not a generative model.
This single tagging step saves enormous time. You stop asking one tool to do everything and start assigning work to the technique that suits it.
Stage 2: Keyframe generation
For every beat tagged key, produce a still image first. Image models are cheaper, faster, and far easier to iterate than video models, and you can judge composition, lighting, and wardrobe in a fraction of a second. Generate three to five candidates per beat, pick one, and save the prompt that produced it. If your image tool supports consistent characters or reference images, use that feature here — it is the cheapest continuity insurance you will ever buy.
Stage 3: Image-to-video animation
Animate the approved keyframes. Keep prompts focused on movement, not description: the image already describes the scene, so a prompt that re-describes the character's clothing only competes with the reference. Describe camera motion, subject motion, and the emotional register. A prompt like slow dolly in, subject turns head slightly, soft window light shifting, cinematic, no text does more than three paragraphs of scene description.
Stage 4: Text-to-video inserts
Generate the texture beats from scratch. This is where you can be adventurous: swarms of particles, macro shots of liquid, aerial cityscapes, abstract light. Because nothing has to match, you can accept the first strong result and move on. This stage is also where you build your transition library — a folder of five-second looping clips that will save you on every future project.
Stage 5: Continuity pass
Line up every generated clip in your editor, mute the audio, and watch it once at double speed. You are looking for drift: skin tone shifts, wardrobe changes, lighting direction flips, a background that rearranges itself. Fix the worst offenders by regenerating or by inserting a cutaway. Do not try to fix everything — audiences forgive small inconsistencies and notice only the glaring ones.
Stage 6: Assembly, sound, and grade
Cut to the rhythm of your music or narration, layer sound effects under every cut to smooth transitions, and apply a light grade so all clips share a colour identity. A subtle film grain or halation layer over the entire timeline does more for perceived coherence than any single regenerated shot.
Prompting for motion, not just appearance
The single biggest skill upgrade in AI video is learning to describe movement. Appearance prompts plateau quickly; motion prompts are where results separate.
Camera language
Borrow the vocabulary of a real camera crew: dolly in, dolly out, truck left, crane up, orbit, handheld follow, static tripod, rack focus, whip pan. Most video models respond to these terms more predictably than to emotional adjectives. Combine exactly one camera move with one subject action per shot. Two camera moves in one prompt usually means the model picks neither.
Subject action and physics
Describe what the subject does and how the world reacts. She lifts the cup, steam curls upward, fabric settles gives the model a chain of physical consequences to render. Words like slowly, gently, and gradually reduce the chance of a jerky, hyperactive result, which is the most common failure mode of short generated clips.
What to leave out
Avoid stacking style words. Cinematic, 8K, photorealistic, filmic, masterpiece at once is noise, not precision. Avoid requesting on-screen text — most models mangle lettering, and typography belongs in your editor where you control the font. Avoid describing things that contradict your reference image when using image-to-video; the model will try to satisfy both and produce mush.
Keeping characters and locations consistent
The hardest problem in generative video is the same face in two different shots. There are four practical levers, and you should use at least two of them together:
- Anchor with a still. Generate one canonical image of the character or location, then use it as the reference for every related shot. Consistency starts with a fixed visual source of truth.
- Freeze the wording. Reuse the exact same descriptor sentences across prompts. Change only the camera and action lines. Small wording variations produce surprisingly large visual variations.
- Control the frame. Prefer medium and wide shots over tight close-ups, and keep faces partly in profile or turned away. Faces break first, so give them fewer pixels and fewer seconds of screen time.
- Hide the seams. Use cutaways, hands, over-the-shoulder framing, and reaction shots to bridge moments where consistency is weakest. This is standard film grammar, and it works for the same reason it works in live action.
Locations are easier than characters. Keep a reference image for each set, describe the light direction explicitly, and avoid camera moves that reveal areas the model has never been shown.
Where AI video projects usually fail
Most disappointing results trace back to a handful of recurring errors:
- Generating before writing. Without a beat sheet you produce attractive clips that never assemble into a story.
- Asking one technique to do everything. Text-to-video cannot hold a product label steady; image-to-video cannot invent an aerial cityscape. Match the technique to the job.
- Chasing the perfect clip. The tenth regeneration rarely beats the third by enough to justify the time. Move on and fix it in the edit.
- Ignoring sound. Silent AI clips feel synthetic. Footsteps, room tone, cloth movement, and a music bed change perception immediately.
- Cutting on motion only. Cut on emotion instead. A cut that lands on the beat of a line or a look reads as intentional; a cut that lands merely when a camera move ends reads as mechanical.
- Skipping the grade. Clips from different models have different contrast and colour science. A shared grade is the cheapest coherence you can buy.
Planning time, compute, and iterations realistically
Estimate your project by shot, not by minute. A realistic first pass on a 60-second piece with 15 beats looks like this: one hour for script and beat sheet, two hours for keyframe exploration, two to three hours for animation attempts including failures, one hour for text-to-video inserts, and two to four hours for editing, sound, and grade. Experienced creators compress this considerably, but the ratio stays similar — planning and editing dominate, generation is a small slice.
Budget iterations per shot rather than per project. Three to five attempts for a key shot, one or two for texture. If a shot resists after five attempts, the prompt is probably asking for something the model cannot do; simplify the request or change technique rather than trying harder.
Choosing a model category for each shot
Rather than memorising brand names, learn the categories and match them to your beats.
| Shot type | Best technique | Why |
|---|---|---|
| Hero product shot | Image-to-video, minimal motion | Reference must stay legible |
| Character close-up | Image-to-video from an approved still | Face consistency is the bottleneck |
| Establishing landscape | Text-to-video | No continuity burden, freedom to explore |
| Abstract transition | Text-to-video, short | Cheap, forgiving, reusable |
| Dialogue scene | Edit-driven with reaction shots | Generation cannot carry performance alone |
| Text and UI overlays | Standard editing software | Typography and timing need control |
A practical rule: if a viewer would notice a change, anchor it with a reference image. If they would not, generate freely.
FAQ
How long should each generated clip be?
Two to five seconds is the sweet spot for most models. Assemble longer sequences in the edit from several short clips rather than requesting one long take.
Do I need a dedicated image model if I am only making video?
It helps enormously. Stills are faster and cheaper to iterate, and they give image-to-video models a stable anchor. Skipping the still stage usually costs more time overall.
Why does my character's face change between shots?
Because each generation starts from a slightly different statistical starting point. Anchor every shot to the same reference image, reuse identical wording, favour wider framing, and cover the weakest moments with cutaways.
Can I fix a bad clip with settings instead of regenerating?
Sometimes. Reducing motion strength or shortening the clip often rescues a warped result. If two attempts at reduced motion still fail, change the underlying keyframe.
How do I stop generated video from looking synthetic?
Sound design, a shared grade, and edit rhythm do more than any prompt. Add room tone under every scene, cut on emotional beats, and apply one consistent look across the timeline.
What is the fastest way to get better at this?
Build a five-second transition library. Every project you make will need inserts and transitions, and having a personal library removes the most repetitive part of the work.
Final checklist
Before you call a project finished, confirm these seven points:
- Every key shot is anchored to an approved reference image.
- No clip exceeds five seconds unless it genuinely earns the extra length.
- Camera and subject motion are described explicitly in each prompt.
- Character and location descriptions are worded identically across related prompts.
- Sound effects, room tone, and music are present under every cut.
- A single grade unifies clips from different model categories.
- You watched the piece once with sound off and once with eyes closed — the first pass catches visual drift, the second catches pacing problems.
Generative video rewards process far more than it rewards luck. Text-to-video gives you reach, image-to-video gives you control, and editing gives you the thing neither can produce on its own: a sequence that feels intentional. Build the beat sheet, tag the shots, anchor what matters, and let the models handle only the parts of the job they are actually good at.




