Why character identity drifts between AI-generated shots
Ask anyone who has produced a narrative sequence with generative video tools what the hardest part was, and you will rarely hear about motion realism or render time. The answer is almost always the same: keeping the same person looking like the same person from shot one to shot twenty.
The reason is structural, not cosmetic. Most image and video models sample each frame independently from a probability distribution shaped by your prompt. Identity is not stored anywhere inside the model as a persistent object. It is implied by tokens, and tokens are fragile. Change the word order, swap a lighting adjective, move from a close-up to a wide shot, switch aspect ratio from 16:9 to 9:16, or update the model version, and the sampling path shifts. Facial geometry follows the shift. The result looks plausible in isolation and wrong in sequence.
The failure modes are predictable once you know what to look for:
- Slow face morphing. Each shot is close, but by shot eight the jawline has narrowed and the eyebrows have moved.
- The sibling problem. Characters remain the same type, same age, same palette, but they are clearly different people.
- Wardrobe and prop drift. A jacket changes colour, a necklace disappears, a scar moves to the other cheek.
- Age sliding. Children become teenagers over a sequence; middle-aged characters lose ten years between a wide and a close-up.
- Skin tone shift from lighting. A warm sunset prompt turns a character warmer than the reference, and the audience reads it as a different person under different light rather than a continuity error.
Every one of these problems is a conditioning problem, and every one of them gets dramatically easier when the model is given more than one image to anchor on. That is the practical value of multi-image fusion: instead of describing a person in words, you show the model several views of that person and let the conditioning stack hold the identity steady while the prompt handles everything else.
What multi-image fusion actually does
The word fusion is used loosely across the ecosystem, so it helps to separate the concept from the branding. At the technical level, multi-image fusion means conditioning a generative model on several reference images simultaneously, rather than on text alone or on a single starting frame.
Different tools implement this differently, but the underlying mechanisms fall into a handful of families:
- Image prompt blending. Two to four images are encoded and blended into a style-and-identity signal that steers generation.
- Adapter-based subject conditioning. A lightweight adapter injects identity features extracted from reference photos into the diffusion or transformer process without retraining the base model.
- Reference-only control layers. A control map derived from a reference image locks pose, silhouette, or facial landmarks so the generator cannot improvise them away.
- Subject reference inputs in video models. The clip generator accepts one or more stills of a character and treats them as persistent subjects across frames.
- Lightweight personalisation training. A small custom weight set is trained on a curated image set so the identity becomes part of the model rather than part of the prompt.
Where each signal lands in the pipeline
It helps to think of generation as four competing inputs. The text prompt carries intent: what is happening, where, with what camera and light. The reference images carry identity: who this is, down to bone structure and hairline. The control maps carry geometry: pose, framing, depth. The seed and sampler settings carry texture and noise character.
When these inputs agree, you get a clean shot. When they conflict, the model averages them, and averaging is what produces the uncanny almost-the-same-face effect. If your references show a character in three different lighting setups, the model splits the difference and produces a fourth, inconsistent lighting on the face. If your reference is a three-quarter profile and your prompt asks for a straight-on close-up, the model guesses at features it cannot see. Reference quality beats reference quantity every time.
When fusion beats fine-tuning
Training a small personalisation set can produce excellent identity retention, but it has real costs: you need a clean dataset, compute time, and a workflow for retraining when the character changes costume or age. For most projects, multi-image fusion is the better default because it is fast, reversible, and easy to iterate. Fine-tuning becomes worth it when you are producing dozens of sequences with the same cast, or when you need the identity to survive aggressive stylisation.
Building a character reference kit that actually works
Most consistency failures trace back to the reference set, not the model. A good kit is small, deliberate, and internally consistent.
Aim for five to twelve images per principal character, with this coverage:
- One neutral, front-facing portrait in even, soft light. This is your ground truth.
- One true profile and one three-quarter view, both in the same lighting as the portrait.
- One full-body shot showing proportions, posture, and default wardrobe.
- Two or three expression variants: relaxed, smiling, serious. Expressions matter more than angles for dialogue scenes.
- One shot with the character's hands visible, if hands will appear on camera.
Technical hygiene matters too. Keep references at similar resolution and aspect ratio. Strip heavy colour grading, since a strongly stylised reference will push the entire sequence into that grade. Avoid sunglasses, hats that hide the hairline, motion blur, and low-light noise. If two characters appear in the same reference, crop them apart before using the image for conditioning; mixed references are a common source of blended faces.
Finally, write the character bible. Two or three sentences, eighty to one hundred and twenty words, describing age, build, hair, distinguishing features, and default wardrobe. You will paste this text verbatim into every prompt, and its main job is to stop you from paraphrasing yourself between shots.
A practical fusion workflow from script to final render
This is the sequence that holds up across projects, whether you are making a two-minute brand film or a twelve-shot narrative short.
Step 1: Lock the character bible and shot list first
Write the character description and the shot list before generating anything. Decide framing, camera movement, and duration per shot. Continuity problems are often storyboard problems in disguise: if you have not decided that shot four is a close-up, you cannot plan the reference you will need for it.
Step 2: Generate a canonical character sheet
Use your strongest still-image model to produce a clean character sheet: front, three-quarter, and profile in consistent lighting. Iterate here until you have a face you genuinely want for the whole project. Everything downstream inherits the decisions you make at this stage. This is the one place where spending extra generation attempts pays for itself.
Step 3: Curate the conditioning set
From the sheet plus any external references, select three to six images for fusion. Fewer, better-matched images outperform a large mixed set. If your tool supports weighting individual references, give the neutral portrait the highest weight and let the angle variants support it.
Step 4: Generate keyframes shot by shot
Generate still keyframes before touching video. Working in stills is faster, cheaper to iterate, and lets you judge identity at full resolution. Reuse the same reference set, the same identity block in the prompt, and a narrow family of seeds across all shots in a scene.
Step 5: Animate with minimal prompting
When you move a keyframe into an image-to-video model, resist the urge to re-describe the face. The keyframe already contains the identity. A short motion prompt, one camera instruction, and a low motion strength will preserve detail. Long, florid motion prompts invite the model to reinterpret the subject.
Step 6: Repair locally instead of regenerating globally
When one shot drifts, fix that shot. Re-fuse with an additional reference, tighten the motion prompt, or shorten the clip. Regenerating an entire sequence to fix one face is the single most common way projects burn time.
Step 7: Assemble, match, and review
Bring clips into your editor, apply a shared grade, and watch the sequence at normal speed with sound. Continuity errors that are invisible in frame-by-frame review become obvious in motion.
Keyframe locking and shot-to-shot continuity
Keyframe locking is the habit that separates hobby output from production output. The idea is simple: for every shot, generate a specific still that the animation step must respect, and where the tool allows it, define both the first and last frame.
First-and-last-frame control is especially useful for dialogue and coverage. If you generate a medium shot and a close-up of the same moment with the same reference set, then animate between them, you get a cut that reads as a genuine camera change rather than two unrelated generations. The identity stays anchored because both endpoints were conditioned identically.
Three practical rules make keyframe locking work:
- One keyframe, one shot. Never reuse a keyframe across shots with different framing; the crop changes what the model sees.
- Document the seed family. Log the seed, reference set, and prompt block for every shot so a reshoot can reproduce the original conditions.
- Extend rather than regenerate. When you need a longer shot, extend the existing clip instead of generating a new one from text, so the identity carries forward frame to frame.
Prompt patterns that keep a face intact
Prompt structure is the cheapest consistency lever you have. Build a fixed identity block and treat it as immutable text.
A reliable pattern looks like this: identity block, then wardrobe, then action, then environment, then camera, then lighting, then style. For example, in a scene where a character walks through a market, the prompt would pair the unchanging description of her appearance with a changing action and setting. The identity sentence never changes. Only the second half of the prompt moves.
A few rules that consistently improve retention:
- Do not paraphrase the identity block between shots. Copy and paste it, character for character.
- Do not re-describe facial features in detail when you are already supplying reference images. Redundant description competes with the references.
- Keep negative prompts stable across a sequence, and include terms that suppress the problems you actually see, such as duplicate faces, extra limbs, or face morphing.
- Keep aspect ratio and resolution constant across a scene. Changing them mid-sequence forces the model to re-frame the subject, which invites drift.
- Put camera and lighting instructions late in the prompt. Early tokens tend to carry more weight, and you want that weight on identity.
Choosing a model and platform for consistency work
The tools matter, but the criteria matter more. When you evaluate an image or video model for a character-driven project, score it against this list:
| Criterion | What to check | Why it matters |
|---|---|---|
| Reference support | How many images can be supplied at once, and can they be weighted? | More and better references mean tighter identity |
| First and last frame control | Can you pin both endpoints of a shot? | Enables true coverage and reliable cuts |
| Seed control | Are seeds reproducible across sessions? | Continuity across weeks of production |
| Motion strength control | Can you dial motion down without killing it? | Prevents the model from reinterpreting the face |
| Output resolution and upscaling | Native resolution and available upscalers | Facial detail survives scrutiny |
| Batch behaviour | Can you queue variations cheaply? | Iteration speed on keyframes |
| Commercial licensing | Clear terms for your use case | Avoids a legal surprise at delivery |
| Predictable compute budget | How costs scale with attempts and clip length | Lets you plan a shoot without surprises |
A useful workflow is hybrid: use one model for the character sheet and keyframes, where identity quality is paramount, and a second model for animation, where motion quality is paramount. As long as the keyframe is strong and the animation prompt is short, identity transfers well. What you should avoid is switching the still-image model mid-project, because each model has its own facial prior and the sequence will subtly shift.
Common mistakes and how to fix them
Too many references. Eight mismatched images are worse than four matched ones. Cut the set down until every image agrees on lighting, age, and grooming.
Mixed-lighting references. If your set contains one golden-hour shot and one studio shot, the model will produce a third lighting on the face. Normalise before you condition.
Re-describing the face in prose. Elaborate facial descriptions fight with the reference images and push the model toward a generic idealised face. Keep the identity block factual and short.
Changing the model mid-project. Even a version bump can shift facial priors. Lock your model and settings, or plan a deliberate migration test on a few shots before committing.
Over-long motion prompts. Every extra clause is another chance for the model to reinterpret your subject. Describe motion, not appearance.
No shot log. Without a record of seeds, references, and prompts, a single reshoot forces you to regenerate a scene. Keep a simple spreadsheet with one row per shot.
Global regeneration for local problems. Fix the shot, not the sequence. Targeted repair is almost always faster and cheaper.
Ignoring the edit. Continuity is judged in motion, with sound. Always review the assembled sequence before deciding a shot is broken.
A pre-export quality checklist
Before delivery, run the sequence through this checklist at full resolution:
- Identity. Pause on every shot and compare against the character sheet. Watch for jawline, eyebrow shape, and hairline, which drift first.
- Wardrobe and props. Check colour, logos, jewellery, and any object the character carries.
- Temporal stability. Play the sequence at normal speed and look for face morphing near the mid-point of clips, where drift most often appears.
- Hands and teeth. These are the highest-frequency artefacts and the first thing a client notices.
- Colour and lighting continuity. Confirm that skin tone stays within a narrow band across shots from the same scene.
- Audio sync. If dialogue is present, verify lip movement against the track.
- Resolution and compression. Confirm the export meets delivery specs without softening facial detail.
FAQ
How many reference images should I actually use?
For most models, three to six well-matched images give the best balance of identity strength and lighting consistency. Start with a neutral portrait, one three-quarter view, and one profile from the same lighting setup, then add expression variants only if dialogue shots are drifting.
Is multi-image fusion a replacement for training a custom character model?
No, they solve different problems. Fusion is faster, reversible, and ideal for projects with a handful of sequences and evolving costumes. A trained personalisation set is worth the effort when the same cast appears across many productions or when heavy stylisation is eroding identity.
Why does my character look right in stills but wrong in video?
The animation step reinterprets the frame. Lower the motion strength, shorten the motion prompt, and avoid re-describing the face in the video prompt. The keyframe already carries the identity; the video model only needs to know how the shot moves.
Can I keep consistency across a cut from wide shot to close-up?
Yes, if you generate both keyframes with the same reference set and the same identity block, then animate with first-and-last-frame control. The apparent camera change reads as coverage rather than a new generation.
What do I do when two characters appear in the same shot?
Condition each character separately where the tool allows it, and keep both identity blocks short. Blended faces usually come from references that contain more than one person, so crop your conditioning images to a single subject.
How do I handle a costume change without losing the face?
Keep the same reference set and change only the wardrobe clause in the prompt. If the model starts altering the face along with the clothing, add one reference image of the character in the new outfit and give it a lower weight than the neutral portrait.
Does resolution really affect consistency?
It does. Changing aspect ratio or resolution mid-sequence forces the model to re-frame the subject, which is one of the most common triggers for drift. Pick your delivery format early and stay with it.
How long should a single generated clip be?
Shorter clips drift less. Generate in short segments and join them in the edit, extending existing footage where possible rather than generating fresh clips from text.
The pattern behind all of this is straightforward. Text alone cannot hold a face. A single image helps, but only from one angle and under one light. Multiple, deliberately chosen references give the model enough evidence to keep a character stable while you change everything else around them. Build the reference kit carefully, lock your keyframes, keep your prompts disciplined, and the hardest part of AI video production becomes a repeatable process rather than a lottery.


