What Multimodal Prompting Actually Changes in Video Production
Text-only video generation asks a model to invent everything: the face, the wardrobe, the room, the light, the lens, the pace. You describe a scene in a sentence and hope the result lands somewhere near your mental image. Sometimes it does. Most of the time it lands close enough to be frustrating — the right mood, the wrong jacket; the right camera move, a completely different character.
Adding a reference image changes the nature of the task. Instead of inventing, the model is now matching. You supply the visual anchor — a portrait, a product photo, a location still, a mood board frame — and the text supplies motion, timing, camera behavior, and dramatic intent. The division of labor is the entire point: images are good at identity and texture, words are good at action and sequence, and neither is good at doing the other's job.
That shift has practical consequences. Shot-to-shot consistency stops being luck and becomes a production process. A character keeps the same face and coat across eight shots because you keep feeding the same reference. A location stays recognizable because you lock a single establishing frame and reuse it. Iteration gets cheaper too: when a take fails, you can usually tell whether the image anchor or the text direction was the weak part, which means you fix one variable instead of rewriting your whole approach.
The rest of this guide is a workflow for that approach — how to split information between image and text, how to prepare references, how to write prompts that direct motion, and how to review output like an editor rather than a gambler.
How Image and Text Prompts Divide the Work
Two input channels, two jobs. Confusing them is the single most common reason a generation drifts away from what you imagined.
Text prompts carry intent, motion, and time
A text prompt should answer: what happens, in what order, seen from where, at what pace, and with what emotional temperature. Verbs and camera language do the heavy lifting. "She turns from the window, hesitates, then walks toward the door; slow dolly in, shallow depth of field, late afternoon light" gives the model a sequence and a point of view. Words like "cinematic" or "beautiful" add almost nothing on their own because they describe a judgment, not an instruction.
Reference images carry identity, style, and continuity
A reference image should answer: who or what is on screen, and what does the world look like. Faces, silhouettes, wardrobe, product geometry, color palette, set dressing, and lighting direction all travel better as pixels than as adjectives. When you need the same person in six shots, an image is the only reliable way to hold that identity — no stack of adjectives will reproduce a specific nose, jawline, or jacket cut.
Building a Shot Brief Before You Generate Anything
Most failed generations are really failed planning. Before you open a tool, write a short brief for each shot. It takes three minutes and saves an hour of rerolling.
The six-line shot brief
A workable template: subject (who or what), action (the single motion that matters), setting (where, with one or two defining details), camera (framing plus movement), light and palette (source, direction, mood), and duration or beat length. Six lines, no prose poetry. If you cannot fill in a line, that is the gap the model will fill for you — and it will fill it with something generic.
Deciding what belongs in the image versus the text
Ask a simple question: is this a thing or an event? Things — faces, products, rooms, color schemes — belong in the reference image. Events — turning, walking, revealing, opening, reacting — belong in the text. If you find yourself typing "a woman with green eyes and a red scarf" into every prompt, that is a signal: make a reference image instead and stop repeating yourself. Repetition across prompts is a maintenance cost, and it drifts.
Writing Text Prompts That Direct Motion
Camera language, subject action, and environment
Describe motion in layers: subject motion, camera motion, environmental motion. "She lifts the cup (subject), slow push in on her hands (camera), steam curling upward and rain streaking the window behind her (environment)." This layered phrasing gives the model three separate things to animate, which produces richer footage than one global verb. Keep one dominant motion per shot. Two competing movements in a single clip usually cancel each other into mush.
Be concrete about speed and distance. "Slow" and "slight" behave differently from "fast" and "sweeping," and models respond to those words more reliably than to mood adjectives. If a shot needs to feel calm, say "slow, steady, minimal movement" rather than "peaceful atmosphere."
Negative guidance and failure modes
List what you do not want, in plain terms: no text overlays, no logo morphing, no extra fingers, no sudden crowd, no jump cut, no camera shake. Keep the list short and specific — five to eight items. Long negative lists tend to suppress legitimate detail along with the artifacts. Track which failures recur in your project; those are the ones worth writing down permanently and reusing in every prompt.
Length and structure
Short prompts produce generic results; enormous prompts produce contradictions. A useful target is one clear paragraph: a subject clause, an action clause, a camera clause, and a lighting clause. Save the rest for the shot brief, which is for you, not for the model.
Preparing Reference Images That Survive Generation
Framing, lighting, and cleanup
Use the cleanest image you can find or make: single subject, uncluttered background, even lighting, no watermarks, no heavy filters, no compression artifacts. Crop tight on the subject if identity matters most; keep a wider frame if you also need to convey environment. If a face is half in shadow, expect the model to guess at the shadowed half in every shot. Upscale before you generate rather than trying to repair softness afterward.
Consistency across characters, props, and locations
Build a small asset library: one neutral portrait per character, one or two angles per product, one establishing frame per location, plus a palette reference. Name files clearly, for example character-name-neutral or location-kitchen-wide, so you can reuse them without hunting through folders. Consistency is mostly bookkeeping — the teams that stay consistent are the ones who keep a tidy asset folder and actually open it every session.
How many references to feed at once
One primary reference plus at most one supporting reference is a good default. Two strong, complementary images — a face and a full-body frame — usually beat five images that contradict each other. When references conflict on lighting direction or wardrobe, the model averages them, and averaging produces a character who looks like nobody.
A Step-by-Step Workflow From Script to Finished Sequence
Step 1: Break the script into shots
Turn every script page into a numbered shot list. One action per shot. If a sentence contains "and then," split it. This is the same discipline live-action storyboards use, and it pays off twice: fewer ambiguous prompts and easier editing later.
Step 2: Write the shot brief for each entry
Fill the six lines. Mark which lines will be handled by a reference image and which by text. Anything about identity or world design gets an image; anything about motion or sequence gets words.
Step 3: Build or collect references
Generate clean stills for characters and locations first — you can use a still-image model, a photo shoot, or existing footage frames. Approve them before animating anything. Fixing a character design at this stage costs minutes; fixing it after twenty clips costs a day.
Step 4: Generate one hero shot first
Do not batch-generate the whole sequence. Produce a single representative shot, review it, and only then expand. That one clip tells you whether your lighting language, camera vocabulary, and reference quality are working together.
Step 5: Iterate on one variable at a time
When a take fails, change either the text or the reference, never both. Keep a simple log: prompt, reference used, what was wrong. After a few sessions you will see patterns — maybe your model handles "dolly in" well but ignores "crane up," or maybe your night references are consistently too dark.
Step 6: Assemble, then patch
Cut the approved clips together before polishing individual shots. Editing reveals which shots are genuinely weak and which only felt weak in isolation. Patch only the shots that break the sequence; perfect singles that do not cut together are wasted effort.
How to Choose Tools and Models Without Locking Yourself In
Quality-first versus speed-first
Some generation tools favor photoreal detail and slow, expensive renders; others favor fast iteration with occasional artifacts. For narrative work, start speed-first: you will throw away most early takes anyway. Once the sequence is locked, regenerate key shots on the higher-quality path. Paying for fidelity before the edit is settled is the fastest way to waste a budget.
Control features that actually matter
Prioritize tools that support image conditioning, camera-motion controls, negative prompts, aspect-ratio control, and clip-length options. Secondary features — templates, style presets, stock sound — are convenient but rarely the deciding factor. Also check export options early: resolution, frame rate, and whether you can get a clean file without an overlay.
Avoiding lock-in
Keep your shot briefs, references, and prompts in plain files you own. If your entire project lives inside one tool's interface, switching costs become enormous. A folder of PNGs and a text file of prompts is portable across every generator you will try next.
Common Mistakes and How to Fix Them
Mistake: one giant prompt for a whole scene
Long, novelistic prompts produce one vague clip instead of several precise ones. Fix: split the scene into shots and give each shot a single dominant action.
Mistake: descriptions doing a reference image's job
If your prompt spends 40 words on appearance, the model has less attention for motion. Fix: move appearance into an image and shorten the prompt.
Mistake: inconsistent references
Alternating between two similar portraits produces a character who changes subtly and unsettlingly between shots. Fix: pick one canonical reference per character and stick to it, noting it in your log.
Mistake: accepting the first usable take
The first take that is not broken is rarely the best. Fix: generate three to five variations with small text changes, then choose.
Mistake: ignoring motion blur and shutter language
Some shots need the crispness of a high shutter and some need the smear of a slow one. Fix: describe it — "sharp, minimal motion blur" or "soft blur on moving hands."
Mistake: mismatched aspect ratios
Generating in one ratio and delivering in another crops your composition. Fix: decide the delivery ratio before generating anything and stay in it.
Reviewing and Iterating Like an Editor
Watch each clip three times with a different question in mind. First pass: does the action read? Second pass: does the camera behave the way the brief asked? Third pass: does the identity and lighting match the neighboring shots? Separate these judgments, because a clip can be beautiful and still useless if the character's collar changed.
Build a visual bible as you go: approved character frames, approved location frames, palette swatches, and a list of prompt phrases that worked. This document becomes the real asset of the project. New team members can pick it up, and your own next project starts from a higher baseline instead of zero.
Finally, keep a short list of "known good" settings per shot type — portrait close-up, product rotation, wide establishing shot. Reusing proven combinations beats inventing new prompt grammar for every clip.
FAQ
Do I always need a reference image?
No. For abstract shots, textures, landscapes, and transitions, text alone works well. Use references where identity and continuity matter: characters, products, branded environments, and any shot that must match a neighbor.
How long should a text prompt be?
One paragraph, roughly 30 to 60 words, covering subject, action, camera, and light. If you regularly need more, your shot is probably two shots.
What do I do when the character's face changes between shots?
Return to a single canonical portrait reference, generate at the same aspect ratio and framing distance, and avoid mixing references from different lighting setups. Consistency problems are usually reference problems, not prompt problems.
Can I use photographs or footage as references?
Yes, if you have the rights to them. Clean, well-lit, uncluttered images work best. Screenshots with interface elements or compressed stills from video usually degrade results.
How many variations should I generate per shot?
Three to five with small, deliberate differences. More than that rarely improves the outcome and slows the edit. Once you have a strong take, stop and move on.
What is the biggest time saver?
Writing the six-line shot brief. Planning shots on paper prevents the loop of generating, squinting, and guessing what went wrong.
How do I keep a long project coherent?
Maintain a visual bible and an asset folder with clear file names, review clips in sequence rather than in isolation, and approve character and location design before animating any motion.

