You can generate a gorgeous five-second clip of a character and then completely fail to reproduce that same face in the very next shot. That gap between a beautiful single frame and a believable continuous performance is the single biggest obstacle in AI video production today. Multi-image fusion is the technique that closes most of it: instead of describing a person in words and hoping, you supply several images of the same person and let the model anchor identity in its own representation space. This guide walks through how fusion works, how to prepare references, how to combine it with keyframe control, and how to build a repeatable workflow that survives cuts, camera moves, and wardrobe changes.
Why Character Continuity Breaks in Generative Video
Most video models are conditioned per shot. Each generation begins from a fresh noise field plus whatever prompt and reference material you hand it. Nothing in that process inherently knows the character existed in a previous clip. The moment camera angle, framing, or action changes, the model reinvents the face: a few millimetres of cheekbone drift here, a new nose bridge there, a jaw that suddenly reads narrower.
Motion makes it worse. During fast turns and quick gestures, the model has fewer stable pixels per frame to reason about, so identity features smear. Wardrobe introduces a second axis of drift: a jacket described as "olive" in one prompt and "dark green" in the next produces two different jackets, sometimes two different fabric weights. Style passes add a third: a colour-grade or restyle step that shifts contrast can also shift skin tone and hair darkness.
Avatar-driven approaches solved part of this by locking a face rig onto a driving performance, but they cap expression range and handle full-body blocking, stylised characters, and non-human designs poorly. The more durable answer is to anchor identity inside the generative model itself rather than patching it afterward. That is the job multi-image fusion performs.
What Multi-Image Fusion Actually Does
Latent anchoring across reference frames
Instead of a single portrait, you provide a small set of images of the same character from different angles and expressions. The pipeline encodes each one into an identity embedding, then merges those embeddings into a single anchor that conditions every generated frame. Because that anchor lives outside any individual shot, the character persists through camera moves, cuts, and scene changes. The generation still varies — lighting and pose change as directed — but the underlying identity signal stays constant.
Reference weighting and conflict resolution
Not every reference deserves equal influence. A sharp, neutrally lit frontal shot is usually the strongest identity signal. A dramatic side-lit profile contributes silhouette and pose information but weaker identity data. Systems that let you tag or weight references produce far cleaner results than systems that average everything blindly. Conflict is the real danger: if you include two references with different hairstyles or eyebrow shapes, the merge tends to split the difference and produce a face resembling neither. Prune contradictions before you generate, not after.
Where fusion ends and prompting begins
Fusion locks identity. It does not lock wardrobe, props, performance, or environment. Those still come from prompts, keyframes, and clothing references. A useful mental model: fusion is the skeleton, prompting is the muscle. A precise identity anchor combined with a vague prompt still yields the wrong scene — just with the right face in it.
Building a Reference Pack That Works
A shot coverage checklist
Aim for six to twelve images. A practical minimum set: frontal neutral, three-quarter left, three-quarter right, profile, slight up-angle, slight down-angle, a smiling frame, a serious frame, a mouth-open frame for dialogue, one full-body shot, and two or three frames lit like your target scene. Coverage matters more than volume; twenty near-identical selfies add almost nothing.
Matching lighting and expression
If your film is lit cool and moody, do not build the pack entirely from warm indoor photos. Mixed lighting pushes the model toward a compromise skin tone that matches neither. Keep wardrobe consistent within a pack, and when the story genuinely changes an outfit, build a second pack rather than trying to describe the change in text.
Cleaning and cropping for signal quality
Remove motion blur, heavy beauty filters, compression artifacts, and busy backgrounds. Square crops centred on the head with a little headroom outperform wide full-frame shots. Keep resolution consistent across the pack so one oversized file does not dominate the merge. If you only have low-quality references, expect softness in the output — the anchor can only be as sharp as its inputs.
Keyframe Control: Directing Poses Without Losing the Face
Posing with skeletons and depth passes
Keyframe control lets you specify body position, camera angle, and depth layout for a shot. Combined with an identity anchor, this is where consistency becomes useful rather than merely accurate. You can drive a walk cycle, a turn, a seated conversation, and a fight beat with the same character because the pose comes from the keyframe and the face comes from the anchor. Keep keyframe poses plausible for the build and proportions of the reference pack; forcing a stiff, narrow-shouldered reference into a wide action pose invites warping.
Expression control and micro-motion
Fine expression control — eyebrow lift, mouth shape, blink timing — is what separates lifeless output from performance. Small asymmetries help enormously: a slight head tilt, unequal eyebrow height, a delayed blink. Over-symmetric faces read as artificial. If your tool exposes intensity sliders for expressions, start subtle and increase only where the beat demands it.
Blending keyframes across cuts
Across a cut, keep the anchor constant and change only the keyframe, camera, and lighting description. This is the core rule of continuity. When a scene transition changes the character's emotional state, change expression in the keyframe rather than rebuilding identity from scratch.
Model Orchestration: Routing Shots to the Right Engine
Matching shot type to model strengths
Different video engines excel at different shots. Some handle dialogue close-ups and skin detail beautifully but struggle with fast full-body motion. Others are strong at wide establishing shots and environmental movement but soften faces. A production pipeline that routes close-ups to one engine and action to another keeps quality high — provided the identity anchor travels with the shot.
Carrying identity across engines
Export your anchor as a reusable preset rather than re-uploading images per engine. Keep the reference set, crop rules, and any identity parameters documented so a shot rendered in a second engine still belongs to the same person. Rebuilding a character in each tool from scratch is the most common cause of visible inconsistency across a finished edit.
Handling restyles and slow motion
Restyle passes, upscales, and frame interpolation can all nudge identity. Test these steps on a short clip before committing them to a sequence. Slow motion in particular can expose interpolation artefacts around the eyes and mouth, where the model has the least reliable motion data.
A Practical Workflow: One Character, Three Scenes
- Write a one-paragraph character brief: age range, build, hair, distinguishing features, resting expression, wardrobe.
- Assemble the reference pack following the coverage checklist, then clean and crop every image to a consistent size.
- Build the identity anchor and test it on a neutral talking-head shot. Fix the pack if the face reads wrong before moving on.
- Create a continuity sheet listing wardrobe, props, hair state, and any injuries or changes per scene.
- Block each scene as a shot list: shot size, camera move, action, dialogue, duration.
- Generate keyframes for each shot's start and end pose, then pass them through with the anchor and a scene-specific lighting prompt.
- Review each generated shot against the continuity sheet at quarter speed, watching eyes, jawline, and hands.
- Regenerate only failing shots, adjusting pose intensity or prompt specificity rather than rebuilding the anchor.
- Assemble in the edit and check the whole sequence at normal speed for drift you cannot see shot by shot.
Continuity Bibles, Shot Lists, and Review Loops
What belongs in a continuity bible
A continuity document is unglamorous and saves hours. Include a front-facing and profile still per character, wardrobe notes with colour names, hair state per scene, prop descriptions, and a short list of forbidden variations (asymmetrical fringe, glasses, stubble). Written colour names beat adjectives like "earthy". Include lighting notes per location so a character entering a new space does not change skin tone because the prompt changed.
Review gates that catch drift early
Build two checkpoints into the workflow. The first is immediately after anchor creation: generate three test shots at different angles and compare them side by side. The second is a sequence review before final polish: place all shots in timeline order and watch at normal speed. Drift is often invisible in isolated frames and obvious in motion.
Common Failure Modes and How to Fix Them
Face melt during fast motion
Reduce pose extremity, shorten the shot, or split a fast beat into two clips with a cut between them. Motion blur can be added in post more cheaply than it can be generated accurately around an unstable face.
Wardrobe and colour drift
Name colours precisely and reuse the same wording in every prompt. If a garment matters, add a dedicated clothing reference image. Repeated small wording changes are the usual culprit.
Background bleeding into identity
Strong background patterns can leak into the character anchor, producing odd textures on skin or hair. Crop references tightly and avoid including heavily patterned clothing in the identity pack.
Over-anchoring and the frozen-photo problem
Too many near-duplicate references can lock the face so tightly that expression range collapses and the character looks like a moving photograph. Trim the pack, favour variety over quantity, and let expression sliders do their job.
Evaluating Tools: Decision Criteria
When comparing platforms, score them on criteria that matter for continuity rather than on raw clip quality alone.
| Criterion | What to look for |
|---|---|
| Reference handling | Multi-image input with weighting or tagging, not a single portrait slot |
| Keyframe control | Pose, depth, and camera control that respects identity |
| Consistency across cuts | Reliable identity retention when lighting and angle change |
| Export control | Reusable character presets you can carry between projects |
| Review workflow | Fast iteration on single shots without rebuilding the anchor |
| Cost predictability | Clear pricing for iteration-heavy work, not just final renders |
A tool that scores moderately on cinematic polish but strongly on anchor stability will usually save more time overall, because re-generating shots is the expensive part of the process.
FAQ
How many reference images do I actually need?
Six to twelve well-chosen images usually outperform twenty random ones. Coverage of angles, lighting, and expressions matters more than count.
Can multi-image fusion handle stylised or non-human characters?
Yes, and often better than face-swap approaches. Fusion anchors visual features generically, so creatures, animated designs, and masked characters all benefit, though you should include references at several distances.
Why does my character change clothes between scenes without being asked?
Inconsistent prompt wording and missing wardrobe references. Write the garment description once and reuse it verbatim, or add a clothing image to the anchor set.
Should I use the same character anchor for every project?
Store it as a reusable preset, but expect to rebuild the pack when the project's lighting or wardrobe changes substantially.
How do I fix a character that looks slightly off at every angle?
It is almost always a pack problem rather than a model problem. Remove conflicting references, tighten crops, and normalise resolution before you experiment with settings.
Is keyframe control necessary for consistency?
Not strictly, but it is what turns a consistent face into consistent blocking. Without it you get the right person in unpredictable poses.
How do I judge consistency objectively?
Watch the full sequence at normal speed, then pause on close-ups and compare jawline, eye spacing, and hairline against your reference sheet. If you notice drift only when comparing stills, it is acceptable for most audiences.
Consistency is not a single feature you switch on; it is a discipline built from a clean reference pack, a shared anchor, disciplined prompt wording, and a review loop that catches drift before it reaches the timeline. Start with one character and three scenes, document what worked, and the workflow scales to a full cast without a full rebuild each time.


