Why character consistency is still the hardest part of AI video
Text-to-video and image-to-video models have become remarkably good at single shots. Give a model a well-written prompt and you get a believable face, believable fabric, believable light. Ask for that same character in the next shot and the illusion collapses. The nose broadens, the hairline retreats, a denim jacket turns charcoal, and a scar quietly moves to the other cheek.
This is often called identity drift, and it is not a bug in one specific model. It is a consequence of how diffusion-based generators work: every frame is a fresh denoising process, guided by text and whatever conditioning the model receives. If the only conditioning is a sentence, the model reconstructs "a woman with red hair and a green coat" from scratch each time, and its idea of that woman is slightly different on every run.
The practical cost is real. A 60-second narrative piece may contain 15 to 25 shots. If two or three of them show a visibly different protagonist, viewers notice immediately, even if they cannot articulate why. Continuity is a trust signal.
Prompt engineering alone rarely solves this. It narrows the range of outcomes, producing a consistent age or a consistent palette, but it does not lock a specific identity. To lock identity, the model needs visual evidence: reference images that carry the character's actual pixels. That is the job of multi-image fusion.
What multi-image fusion actually does
Multi-image fusion is the practice of feeding a generation model several reference images at once, typically three to six, and letting it blend their visual information into a new shot. Instead of describing a character, you show the model several angles of that character and let the conditioning layer carry identity while your prompt carries action.
Reference roles: identity, wardrobe, lighting, palette
Not every reference should do the same job. A useful mental model is to assign each image a role:
- Identity anchors: close-up, neutral expression, even lighting, high detail on the face.
- Wardrobe references: a full-body or three-quarter shot showing garment cut, fabric, and colour.
- Silhouette references: a profile or back view that pins down hair volume, posture, and head shape.
- Lighting and palette references: a shot that establishes the mood you want the new scene to inherit.
When you hand the model five near-identical selfies, you get five votes for the same information and zero votes for anything else. Role assignment prevents that.
How the model weighs references
Most fusion-capable pipelines accept references either as a concatenated conditioning set, an image embedding blended per-image, or a face or identity adapter that operates alongside the text prompt. Implementation details vary, but behaviour is fairly consistent: images with clean, front-facing, well-lit faces dominate. Blurry, cluttered, or dramatically lit references contribute mostly noise.
That asymmetry matters. One bad reference can drag the whole generation toward its problems. Curate ruthlessly.
Building a reference pack that holds up
Shot selection
Aim for coverage, not quantity. Six strong references beat twenty average ones. A workable pack looks like this:
- Straight-on close-up, neutral expression, eyes open, no strong shadows.
- Three-quarter view, slight smile, same lighting conditions as image one.
- Profile view, relaxed posture.
- Full-body shot in the character's default wardrobe.
- A dynamic or expressive shot, running or laughing, to teach the model how the face deforms.
- An optional detail shot: hands, a signature accessory, a tattoo.
Consistency of lighting between the first two references matters more than most people expect. If one reference is lit by window light and another by a hard flash, the model receives conflicting signals about skin tone.
Preprocessing checklist
Before references go into a pipeline:
- Crop to the character and remove distracting background objects where possible.
- Normalise resolution. Downscaling a large portrait to a consistent 1024 or 2048 pixels on the long edge keeps embeddings comparable.
- Confirm the face is sharp at 100% zoom. Slight camera shake is invisible in a thumbnail and catastrophic in an embedding.
- Check colour space and white balance. Yellow-cast references produce yellow-cast characters.
- Remove watermarks and text overlays.
- Save as PNG or high-quality JPEG; repeated JPEG re-encoding adds blocking artefacts.
Store the pack as a named unit, for example "Maya_v3", and never mix packs between projects. Versioning reference packs is the single cheapest habit for reproducible results.
A repeatable fusion workflow, step by step
Step one: write a character sheet
Before generating anything, write a one-page sheet: age range, build, hair colour and length, eye colour, skin tone, distinguishing marks, default wardrobe, and two or three personality adjectives. Keep it to facts the camera can see. "Sarcastic" is useful for performance prompts; "green eyes" is useful for identity.
Step two: assemble and label the references
Load three to six images from the pack and label each with its role in your notes. If the interface supports per-image weighting, weight the neutral close-up highest and the expressive shot lowest.
Step three: write the anchor prompt
The anchor prompt should do two things: restate the immutable physical facts and describe the new scene. Structure it as an identity clause, then an action clause, then a camera clause.
Example: "Same woman as the reference images: early thirties, dark brown shoulder-length hair, freckles across the nose, grey wool coat. She steps off a tram into light rain, glancing left. Medium shot, 35mm, shallow depth of field, cool overcast light."
Note what is missing: no re-description of hair texture in three adjectives, no mood words that fight the reference. Redundancy is not harmful, but contradiction is fatal.
Step four: test on the cheapest shot
Never start with the hero shot. Generate a short, low-cost test, two seconds with a simple background, and compare it against the identity anchor side by side. Ask three questions: is it recognisably the same person, is the wardrobe correct, and is the lighting compatible with neighbouring shots?
Step five: lock, propagate, and only then vary
Once a setting produces a faithful result, freeze it: same references, same weights, same seed where supported, same resolution. Change one variable at a time when the scene changes. Directors instinctively want to vary everything at once; in a fusion pipeline that destroys your ability to diagnose what broke.
Step six: build a shot library
Save every approved still and short clip to a project folder indexed by character and scene. When a later shot drifts, you can compare against an approved frame rather than against memory. This also gives you re-conditioning material: a previously approved frame is often the strongest reference you have.
Prompt patterns that preserve identity
Certain phrasings help; others reliably hurt.
Helpful patterns:
- "Same character as the reference images, consistent facial features."
- Explicit, concrete physical facts repeated verbatim across shots.
- Camera and lens language, such as "medium shot, 50mm", to stabilise framing.
- Lighting descriptors that match an existing reference, such as "soft window light from camera left".
Harmful patterns:
- Identity-changing adjectives: older, younger, glamorous, rugged. Each one nudges the face.
- Style words that compete with the reference, like "anime" or "oil painting", unless the whole project uses them.
- Contradictory wardrobe instructions. If the reference wears a grey coat, do not ask for a black jacket in the same shot; change the reference instead.
- Very long negative lists. They consume prompt budget and often introduce artefacts.
A useful rule: describe the scene richly and the character minimally. The references handle the character. Your words handle everything else.
Handling motion, lighting, and camera changes
Identity drift gets worse as the shot gets harder. Three variables cause most of it.
Motion. Fast movement forces the model to interpolate, and interpolation is where faces melt. Keep the first shot of any new scene slow or static, then escalate. If a running shot drifts, generate a mid-motion keyframe as an image, approve it, and use that still as an extra reference.
Lighting. A character established in soft daylight will drift if you suddenly cut to neon night. Bridge the change: generate one transition shot with mixed lighting, approve it, and use it as the lighting reference for the night scene. It also reads better editorially.
Camera. Wide shots reduce facial detail, which paradoxically helps consistency because there is less to get wrong, and hurts it because the model has more freedom in silhouette. For wides, rely on wardrobe and hair silhouette as the identity signal.
Angle changes are the real test. Front-to-profile transitions are the most common failure point, which is exactly why a profile reference belongs in the pack from day one.
Where fusion fits in a production stack
Most teams end up with a layered workflow rather than one tool:
- Concept and character design: image generation for exploration, with no identity requirements yet.
- Reference pack creation: curated stills, cleaned and labelled.
- Shot generation: a fusion-capable video model driven by the pack.
- Consistency repair: targeted regeneration using approved frames as extra references.
- Assembly: editing, sound, colour. Colour grading affects perceived identity more than people expect; a warm grade can make two shots of the same character read as different people.
Keep the reference pack and the character sheet in version control alongside the edit. When someone asks why a shot from week three looks different, you can compare the inputs instead of guessing.
Failure modes and how to fix them
The face is right but the age is wrong. This is usually caused by a reference that includes children or older adults in a multi-person pack, or by age-related words in the prompt. Remove young and old phrasing and re-check the pack for outliers.
Wardrobe keeps changing colour. Colour drift typically comes from references captured under different white balance. Normalise the stills, or add an explicit colour word plus a lighting reference.
Everything looks slightly plastic. Often this means too many references at low resolution, all near-duplicates. Replace three similar images with one sharper close-up and one genuinely different angle.
The character looks right in stills but wrong in motion. Motion blur and compression are eating detail. Slow the action, increase resolution, or shorten the shot and cut around the difficult frames.
Results vary between runs with identical settings. Some pipelines are stochastic even with a fixed seed. Generate three takes and pick one, or reduce the number of simultaneous variables.
One shot in ten is unusable, every time. Track it. If failures cluster around a specific shot type, such as profiles or night exteriors, that is a pack gap rather than bad luck.
Quality control across a full sequence
Build a simple review ritual. Lay five to eight consecutive shots in a timeline at thumbnail size. At that scale, identity errors are obvious because you are comparing shape and colour rather than detail. Flag any shot where the silhouette or skin tone jumps.
Then check three technical markers: hairline position, eye spacing, and garment colour. These three fail first and are cheap to verify.
For longer projects, consider a contact sheet with one frame per shot, displayed in a grid. Contact sheets expose drift that a moving timeline hides, because your eye can compare frames directly rather than remembering them.
Finally, decide your tolerance in advance. Animation and stylised work tolerate more variation than live-action-style realism. A documentary-style character piece tolerates almost none. Writing the threshold down saves arguments in review.
FAQ
How many reference images should I use? Three to six is the practical sweet spot. Fewer than three gives the model too little evidence; more than six usually adds redundancy rather than information, and can dilute the strongest reference.
Do I need a different reference pack for each outfit? Yes for the wardrobe layer, no for the identity core. Keep the identity anchors constant and swap wardrobe references as the story moves. Treat the pack as a stable core plus scene-specific layers.
Can text prompts replace references entirely? For short, stylised pieces, sometimes. For anything with recurring characters, no. Text narrows the distribution; images pin a point inside it.
What resolution should references be? Match what the pipeline expects and keep every image in the pack consistent, typically 1024 to 2048 pixels on the long edge. Mismatched sizes produce uneven weighting.
Why does the character change when I change the background? Background and subject are entangled in the conditioning. Add a clean, plain-background reference of the character so identity information stays separate from scene information.
Is a fixed seed enough for consistency? No. A seed fixes the noise, not the identity conditioning. Change the prompt or the references and the character changes even with the same seed.
How do I fix a single bad shot without regenerating the scene? Generate a still first, approve it, then animate from that still with the pack still attached. Frame-level repair is far cheaper than re-running a whole sequence.
Do I still need prompt engineering if I use fusion? Yes, but its role changes. Prompts stop being a description of the character and become stage direction: what happens, how the camera moves, what the light does. That division of labour is what makes multi-image fusion reliable enough for long-form work.



