Why the input channel decides the output quality
Generative video tools have become good enough that the real bottleneck is no longer rendering power. It is direction. Two creators can open the same model, paste prompts of similar length, and walk away with clips that look like they came from different decades. The difference is rarely the model. It is the way each creator selects an input channel and builds a chain of decisions around it.
There are three channels that matter: text, reference images, and custom control parameters. Text is the fastest way to explore an idea and the worst way to reproduce one. Images lock appearance, palette, and silhouette at the cost of flexibility. Custom parameters — duration, motion strength, camera path, seed, aspect ratio, style weight — turn a lucky accident into a repeatable result.
The common mistake is treating these as competing options. They are not. They are stages. Text finds the shot. An image freezes it. Parameters reproduce it. Everything below is built on that order.
The three input channels and when each one wins
Text-to-video: exploration at full speed
Text is the cheapest experiment you can run. One sentence can produce a usable establishing shot, a mood test, or a transition that would have taken an afternoon to shoot. Use text-first generation for mood exploration, abstract B-roll, environment plates, and testing whether an idea has legs before you invest in it.
Do not use text alone for recurring characters, branded product shots, or anything that must match a clip you already approved. The model has no memory of your last output, so a face described in prose will drift between takes.
A useful habit: write three variants of every idea — a literal version, an atmospheric version, and one deliberately wrong version. The wrong one often teaches you which words the model is actually responding to. Delete the noise, keep the discovery.
Image-to-video: locking identity and look
A keyframe converts guesswork into anchoring. When you supply a still, the model inherits composition, palette, wardrobe, and facial structure. That is the single most reliable way to keep a character recognisable across a sequence.
There are two distinct modes, and mixing them up wastes time. The first is animating an existing frame: subtle parallax, drifting hair, breathing, rain, a slow push-in. The second is using a still as a style and composition reference while generating genuinely new motion. The first is a finish, the second is a starting point. Decide which one the shot needs before you touch a slider.
Image-first work excels at character consistency, product footage, architectural reveals, and stylised illustration where the drawing style matters more than physical realism.
Custom parameters: the reproduction layer
This is where professionals separate from hobbyists. Resolution, aspect ratio, clip length, motion strength, seed, guidance strength, style weight, frame interpolation, and explicit camera instructions all live here. None of them are glamorous. All of them determine whether tomorrow's shot matches today's.
When a shot works, write down every setting alongside the prompt. An unlabelled winning clip is a one-off, not a technique. Build a personal recipe file: prompt, reference image path, model, seed, parameters, and a one-line note about what the result actually looked like. After twenty entries you will start seeing patterns — which parameters cause flicker, which seeds give you stable faces, which motion strengths survive a close-up.
Prompt engineering for motion, not for stills
The six-slot shot prompt
Most weak prompts are weak because they describe a subject and stop. Video needs behaviour. A reliable structure is six slots: subject, action, setting, camera, lighting, look. Keep the whole thing under about sixty words unless the model rewards long context.
Here is a filled example:
A tired lighthouse keeper slowly turns from a rain-streaked window as the storm breaks, inside a stone tower room at first light, slow dolly-in from medium shot to close-up, cold blue daylight against one warm lamp behind him, 35mm film grain and a muted teal palette.
Notice what each slot contributed. The subject gives the model a person. The action gives it a beginning and an end, which is what makes a clip feel like a moment rather than a still that jiggles. The camera slot controls the movement. The lighting slot controls contrast and mood. The look slot keeps the output stylistically consistent with neighbouring shots.
Six slots is not a rule, it is a checklist. When a generation fails, identify which slot was vague. Nine times out of ten it is the action or the camera.
Motion vocabulary worth memorising
Vague words produce vague movement. Learn the vocabulary and use it precisely:
- Camera moves: dolly in, dolly out, truck left, crane up, orbit, arc, whip pan, handheld, static lock-off, slow push.
- Shot sizes: extreme wide, wide, medium, medium close-up, close-up, insert, macro.
- Subject motion: steps forward, turns, reaches, exhales, braces, glances, settles.
- Time behaviour: slow motion, real time, time-lapse, speed ramp, freeze at the end.
- Optics: shallow depth of field, deep focus, anamorphic flare, wide-angle distortion, compression from a long lens.
Pairing a camera instruction with a subject instruction is what makes a shot feel directed instead of generated. "She walks" is a description. "Handheld medium shot follows her as she walks away from camera, shallow focus" is a decision.
Negative prompts and known failure states
Negative prompts are a repair tool, not a personality. Keep them short and targeted: warped hands, extra limbs, flickering, morphing faces, duplicated subjects, on-screen text artefacts, smeared crowds, rubbery physics. If your negative list grows past roughly ten items, the prompt itself is under-specified.
Recognise the classic failures. Faces melt during fast motion because the model lacks a stable identity anchor. Backgrounds pulse because motion strength is too high for a detailed scene. Hands warp when they occupy too much of the frame. Text renders as gibberish unless the model supports typography explicitly. Each of these has a specific fix, and almost none of them are solved by adding adjectives.
Building a character and style bible
Reference images that condition well
A reference image is only as useful as it is clean. The best conditioning frames share a few properties: a single subject, an uncluttered background, even lighting without harsh colour casts, no watermarks or logos, and a consistent set of angles — front, three-quarter, and profile at minimum.
Wardrobe consistency matters more than most people expect. If your reference set includes the character in three different jackets, the model will treat the jacket as a variable and invent a fourth. Lock the silhouette before you generate anything you intend to keep.
Seeds, style tokens, and continuity discipline
Reusing a seed is the simplest continuity trick available. It will not guarantee identical output, but it narrows the randomness enough that lighting and grain direction stay in the same family across shots.
Style tokens — a short phrase like "muted teal palette, 35mm grain, soft window light" — do the same job for colour and texture. Repeat the exact phrasing between shots rather than paraphrasing it. Models are literal about tokens and approximate about meaning.
What the style bible contains
- Character reference images and their file names
- Wardrobe and prop definitions with colour values
- The exact style token string used in every prompt
- Preferred lens language and shot-size conventions
- Seed and parameter settings for approved shots
- A list of banned elements: logos, specific celebrity likenesses, unwanted text
A style bible sounds bureaucratic until the first time you need to regenerate one shot in a sequence of thirty. Then it saves an entire afternoon.
Choosing the right model for the shot
Photoreal and cinematic work
When you need believable skin, believable light, and believable physics, you want a model tuned for realism. Runway, Sora, Veo, and Kling all compete in this space, and each has a personality: some favour dramatic camera movement, others favour stillness and detail. Test the same six-slot prompt across two or three of them before committing an entire project to one.
The practical criterion is not which model looks best in a highlight reel. It is which model gives you the fewest unusable frames per minute of footage. Repair time is the hidden cost of every pipeline.
Stylised, animated, and illustrative work
Anime, motion-graphics, and painterly looks often come out better from models that are less obsessed with photorealism. This is where image-first conditioning shines: generate a strong illustration as a keyframe, then animate it, so the art direction is decided before motion enters the equation.
For anything with a strong house style — flat vector, watercolour, comic ink — keep a style reference image in addition to text tokens. Images carry texture in a way words cannot.
Speed-first iteration and social formats
Short-form vertical video rewards volume and speed over pixel-perfect polish. Choose a fast model, generate many short clips, and accept a lower hit rate. Nine mediocre clips and one great one is a perfectly rational output for a thirty-second social edit.
Batch your work. Write ten prompts in one sitting, generate them back to back, then review as a group. Switching between writing and reviewing for every single clip destroys momentum and produces fewer usable results.
Open-source and local pipelines
If you need full control, reproducible parameters, or offline operation, open-weight video models running through a node-based interface give you the deepest parameter access available. The trade-off is setup time, hardware requirements, and a steeper learning curve. This path suits studios with repeated workflows and stable subject matter more than one-off projects.
A hybrid approach works well: use hosted models for exploration, then move approved looks into a local pipeline where every parameter is documented and repeatable.
A repeatable workflow from script to delivery
Step 1 — shot list and storyboard
Before generating anything, write the shot list in words. One line per shot: size, subject, action, approximate duration. This document is your defence against the temptation to generate randomly and hope a story appears in the edit.
Step 2 — keyframe generation
Generate still frames first. Stills are cheap and fast to iterate. When a frame looks right, it becomes the anchor for the shot. Building a project out of approved frames rather than approved clips reduces rework dramatically.
Step 3 — the motion pass
Animate one shot at a time with the six-slot prompt plus the keyframe. Keep motion strength modest for close-ups and dialogue, higher for landscapes and action. Generate two or three takes per shot and choose immediately rather than hoarding variants.
Step 4 — selects, upscaling, and assembly
Mark selects with a consistent naming scheme so the edit imports in order. Upscale only the clips that survive the first cut. Then assemble in a standard editor, trimming to rhythm rather than to full clip length. AI clips rarely earn their entire duration.
Step 5 — sound, grade, and delivery
Sound is where AI video most often gives itself away. Add room tone, foley, and a consistent music bed. A light grade that matches contrast and colour across shots can make footage from three different models feel like one camera. Add captions for silent viewing, export at platform-appropriate bitrates, and check the result on a phone screen before publishing.
Keeping multi-shot sequences coherent
Coherence is a system property, not a model feature. Four things carry most of the load: a fixed style token string, a locked character reference set, consistent lighting direction, and a shot-size rhythm that avoids jumping from extreme wide to extreme close-up without a reason.
When a sequence feels wrong but you cannot say why, check continuity of light direction first. Two shots with daylight coming from opposite sides will read as a mistake even to viewers who cannot articulate it.
The second most common culprit is inconsistent motion energy. A locked-off shot next to a whipping handheld shot creates a jump in energy that no grade can hide. Group shots by energy level during assembly.
Common mistakes and how to fix them
- Prompting a subject instead of a moment. Add a beginning and an end to the action. Movement needs a destination.
- Overspecifying. Long prompts with twelve adjectives dilute the important instructions. Cut to six slots and stop.
- Ignoring aspect ratio early. Generating widescreen footage for a vertical platform wastes detail. Decide the frame before generating.
- Hoarding takes. Fifty variants is not a library, it is a decision you are avoiding. Keep two, delete the rest.
- Skipping the keyframe. If a shot must match an approved look, generate the still first. Always.
- Trusting one model with everything. Different shots need different strengths. A single-tool pipeline is a single point of failure.
- Leaving sound for the end and then rushing it. Budget as much time for sound design as for the motion pass.
A pre-publish quality checklist
Run this before export, every time: Are hands and faces stable across cuts? Does light direction stay consistent? Do colours match between shots? Is motion smooth at normal playback speed, not just paused? Are there any unwanted on-screen text artefacts? Does the audio sit at a consistent level with no clipping? Do captions render correctly in the safe area? Does the first two seconds earn attention without a title card explaining everything?
If a shot fails more than two of these, regenerate rather than repair. In AI video, regeneration is usually faster than fixing.
FAQ
How many words should an AI video prompt be?
Most shots work best between twenty and sixty words. Below fifteen words you are usually missing an action or camera instruction. Above a hundred you are competing with yourself, and the model will weight instructions unpredictably. Expand only when the model explicitly rewards long, structured prompts.
Do I need reference images if my text prompts are detailed?
Only if consistency matters. For a one-off abstract clip, text is enough. For a character appearing in six shots, a reference image set will save you more time than any prompt refinement.
Why do my characters change appearance between shots?
Because text descriptions are approximations, not specifications. Fix it with a locked reference set, a reused seed, and an identical style token string. If drift persists, generate keyframes first and animate from them instead of generating from text.
Is it better to generate long clips or assemble many short ones?
Shorter clips assembled in an editor give you far more control over pacing and are easier to regenerate when one shot fails. Long single generations are convenient but rigid — a flaw at second twelve means starting over.
How do I stop footage from looking like AI?
Three fixes cover most cases: reduce motion strength, add grain and a slight grade to unify texture, and design sound properly. Sound is the fastest credibility upgrade available, and it is the step most creators skip.
Can I mix output from several different video models in one project?
Yes, and many professional sequences do. The unifiers are a shared style token string, a single grade pass, and consistent audio treatment. Vary the model per shot type, not per shot within the same scene.
How much footage should I generate for a finished minute?
Budget roughly five to eight times your final runtime in raw clips for a polished piece, and less for social content where a rougher texture is acceptable. Plan for the ratio rather than being surprised by it.
Where to go next
Pick one shot type — a character close-up, a landscape reveal, or a product turn — and build a complete, documented recipe for it: prompt, keyframe, parameters, seed, and a note on what worked. Then repeat the exercise with a second shot type. After a handful of documented recipes you will have something more valuable than a folder of clips: a method that produces consistent results on demand.
The models will keep changing. The order of operations — text to explore, image to lock, parameters to reproduce — will not.



