Why identity drift is the hardest problem in AI video
A character walks into frame in shot one and looks exactly right. Two shots later the jaw has widened a little, the jacket has shifted from charcoal to blue-grey, and the eyes sit a few millimetres further apart. Nothing is catastrophically wrong, yet the illusion collapses. Human beings read faces with extraordinary precision, and a face that is ninety-five percent correct reads as a different person.
This is the real bottleneck in serialized AI video work. The difficult problem is not generating one beautiful frame — modern models do that easily. The difficult problem is generating the same person across thirty frames, four camera angles, two lighting setups, and a costume change. Early reference conditioning, where you supply a single portrait, solved part of it. Multi-image fusion solved much more, because it gives the model several simultaneous views of the same identity and lets it triangulate what must stay constant.
The practical consequence is uncomfortable for people who love prompt craft: character consistency is a data problem before it is a prompting problem. A weak reference pack cannot be rescued by clever wording. A strong reference pack will carry you through ordinary prompting. Everything below is about building the pack, wiring the pipeline, and checking the result before you commit to a full render.
What multi-image fusion actually does
Fusion is not a single feature you can switch on. It is a conditioning strategy. Instead of describing a person in text, you hand the model a small set of images and let its encoder extract a dense identity representation. When you feed several images, the system compares them, discards what changes between them, and keeps what repeats.
In practice this happens at two levels. First, a vision encoder converts each reference into an embedding. Second, an attention or adapter mechanism injects those embeddings into the generation process, usually weighted by how much identity fidelity you want versus how much freedom the scene needs. Some pipelines — ComfyUI graphs built around IP-Adapter style nodes, for example — let you blend multiple embeddings with individual weights. Others hide the mechanics behind a single "character reference" slot.
The three layers every reference pack encodes
A useful mental model is that your images communicate three different things, and you should control them separately:
- Identity layer: bone structure, eye spacing, nose shape, skin tone, age markers. This is what must never change.
- Presentation layer: hairstyle, wardrobe, accessories, makeup. This may change between scenes, but only when you intend it.
- Capture layer: lens, lighting direction, colour grade, contrast. This is the layer most people accidentally bake in, and it fights every new scene you generate.
If all five of your reference images were shot with a warm key light and a shallow depth of field, the model may treat that lighting as part of the person. Then your night scene looks like a badly graded day scene.
More images is not automatically better
Adding references has diminishing returns and eventually becomes harmful. Beyond roughly six to ten strong images, you start feeding the encoder contradictions: different hair lengths, inconsistent stubble, a slightly different nose after a weight change. The model averages the noise, and you get a slightly generic face that resembles nobody in particular. Curate aggressively. A tight set of eight consistent images beats forty scraped ones every time.
Building a reference pack that survives every shot
Treat the reference pack as a production asset, not a folder of screenshots. The goal is coverage: enough angles and expressions that the model can reconstruct your character at an angle you never supplied.
The core set
For a speaking or recurring character, aim for these six to eight frames:
- Neutral headshot, straight on. Eyes to camera, relaxed expression, even lighting. This is your anchor.
- Three-quarter view. The single most valuable angle for video work, because most dialogue coverage sits here.
- Profile. Prevents the nose and jaw from melting during turns.
- Full body. Establishes proportions, height cues, and silhouette.
- Expressive variant. Laughing or mid-speech, so the model learns how the face deforms.
- Wardrobe sheet. Costume pieces laid out or worn, described identically in every prompt.
- Environment test. The character lit like the actual scene, to check how the grade behaves.
Technical rules that quietly matter
- Resolution: at least 1024 pixels on the short edge, ideally higher. Soft, compressed references produce soft, drifting faces.
- One subject per image. Extra people confuse the encoder, which may fuse unwanted features.
- Vary the background. Identical backgrounds encourage the model to memorise the room instead of the person.
- No heavy filters or beauty retouching. Over-smoothed skin gives you a plastic character that never quite matches the scene.
- Sharp eyes. Eye clarity is the strongest single predictor of whether a cloned character reads as convincing.
Name and version everything
Use a predictable naming scheme such as char-elena-v03-neutral.png. Keep a short manifest listing which images are in the active pack and when each was added. When a character suddenly starts looking wrong after a pipeline update, the first question is always which references were in play. Versioning turns a two-hour debugging session into a two-minute rollback.
A step-by-step multi-shot workflow
Here is a workflow that holds up across a full scene rather than a single clip.
Step 1: Write the identity bible
Before generating anything, write one paragraph describing the invariants: age, build, hair, skin tone, distinguishing marks, wardrobe defaults. This becomes the fixed block you paste into every prompt. Ambiguity here guarantees drift later.
Step 2: Lock the hero frame
Use an image model — Flux, Midjourney, or a Stable Diffusion setup with a tuned adapter — to generate stills of your character. Iterate until one frame is genuinely right. That frame becomes the visual contract for the entire project, and every later shot is measured against it.
Step 3: Generate keyframes before video
Generate a still for each shot you need, conditioned on your reference pack. Reviewing stills is fast and cheap; reviewing video is slow. A storyboard of six approved stills costs far less time than six rejected video clips.
Step 4: Animate with reference conditioning switched on
Move each approved still into an image-to-video model — Runway, Kling, Luma, Sora, or a comparable system — and keep the character reference active during animation. Subtle motion works better than dramatic motion: a slow push-in with a head turn preserves identity far better than a sprint through a crowd.
Step 5: Run a consistency pass, then assemble
Check every clip against the hero frame before editing. Then cut the sequence together and watch it end to end. Drift that is invisible shot by shot becomes glaring at the cut point, because the eye compares the outgoing and incoming faces directly.
Step 6: Upscale and grade uniformly
Apply the same upscaler and the same colour pipeline to every clip. Divergent post-processing is one of the most underrated causes of "my character changed" — the identity is fine, but the two clips no longer share a grade.
Prompt tactics that lock identity in place
Once the reference pack is solid, prompting becomes about restraint.
Separate fixed and variable text
Build prompts from two blocks: a character block that is copied verbatim into every prompt, and a scene block that changes. Rewriting the character description in fresh words for each shot is one of the fastest ways to drift, because synonyms carry different associations. "Slim woman with sharp features" and "slender woman with angular face" do not condition the same way.
Describe less than you think
The reference images already carry the identity. Your text should carry the intent: camera, action, lighting, mood. Long adjective stacks push the model toward its own interpretation of a generic face and away from your reference.
Use negative constraints
Explicit exclusions help more than most people expect. Typical entries include: no change to hairstyle, no change to eye colour, no age shift, no facial hair, no generic beauty retouching, no wardrobe change, no extra characters. Keep the list short and specific; bloated negatives can destabilise composition.
Repeat prop and wardrobe tokens verbatim
If your character carries a red canvas satchel, every prompt says red canvas satchel. The moment one prompt says shoulder bag, the model happily invents a different one — and the audience notices the prop change even if they cannot articulate the face difference.
Choosing the right generation route
There is no single best approach. Match the route to the number of shots and the tolerance for iteration.
Reference-conditioned text to video
Fast and flexible. Good for mood pieces, montages, and short social clips where you need a character to feel consistent but not survive a close-up continuity check. Weakest at maintaining fine facial detail across many shots.
Keyframe-anchored image to video
The workhorse for narrative work. You generate stills with a strong character-conditioned image model, approve them, then animate. This puts your review gate before the expensive step and keeps identity anchored to something you have already approved.
Adapter or lightweight fine-tune
When a character appears in dozens of shots across multiple episodes, training a small adapter on your curated pack pays for itself. It encodes identity more rigidly than prompt-level references and survives harsher camera moves. The trade-off is setup time and a new asset to maintain.
| Situation | Recommended route |
|---|---|
| A handful of clips, loose continuity | Reference-conditioned text to video |
| A scene with dialogue and cuts | Keyframe stills, then image to video |
| Recurring character, many episodes | Small trained adapter plus reference pack |
| Rapid client mockups | Text to video with two or three references |
| Extreme angles or action | Keyframe stills with profile and full-body references |
Quality control: catching drift before the full render
Build a gate that runs before you commit to rendering a whole sequence.
- Contact sheet review. Lay the hero frame and every candidate keyframe side by side at thumbnail size. Similarity problems are easier to spot at small scale than full size.
- The squint test. Shrink clips and play them in sequence. If the character reads as two different people when blurred, no amount of sharpening will fix it.
- Overlay comparison. Align the hero frame and the new frame, then toggle between them. Eye line and jaw width discrepancies appear instantly.
- Similarity scoring. Embedding-based face similarity scores are useful as a fast sort, not as a verdict. Use them to rank candidates, then decide with your eyes.
- Cut-point checks. Review every transition, not every clip. That is where the audience compares faces.
Common mistakes that break consistency
- Mixing references from different eras of a character. A pack assembled over months often contains contradictory hair lengths and weights.
- Over-lit beauty references. Studio-glossy images fight every naturalistic scene and push the grade in the wrong direction.
- Changing seeds, prompts, or settings mid-sequence. Isolate variables; change one thing at a time.
- Ignoring focal length. A character shot at 24mm and 85mm looks subtly different. Match the virtual lens between shots.
- Rendering at maximum resolution too early. Iterate at low resolution, then scale once the shot is approved.
- No version history. Without it, you cannot tell whether a regression came from new references or a model update.
- Relying on text alone. Descriptions approximate a person; images define one.
- Skipping the assembly review. Per-shot approval hides cumulative drift.
Scaling consistency across episodes and teams
Once one character works, the temptation is to treat it as solved. It is not — it needs infrastructure. Keep a shared asset library where every reference pack lives with its manifest and date. Maintain a prompt template with the fixed character block, so nobody rewrites descriptions from memory. Document which model version produced which approved shots; when a pipeline changes, you will want to know exactly what moved. Add a review gate at keyframe stage, not at final render, and give one person ownership of identity continuity. On multi-character projects, define interaction rules too: how the two characters share frame space, and which one holds the visual focus when both are present. That last detail prevents the quiet failure mode where two cloned characters blend features during close interaction.
FAQ
How many reference images do I actually need?
Six to eight well-chosen images cover most needs: neutral, three-quarter, profile, full body, an expression variant, and a lighting test. Below four, drift rises sharply. Above ten, contradictory signals often make the result worse rather than better.
Can I get away with a single reference image?
Sometimes, for short clips with limited camera movement and one lighting setup. The moment you need a profile, a turn, or a different grade, a single image starts to fail. Think of one image as a proof of concept, not a production asset.
Why does my character change after upscaling?
Upscalers hallucinate detail, and faces are where hallucination is most visible. Changes in eye shape or skin texture after a second pass usually mean the upscaler is doing too much. Use a gentler model, apply the same settings across all clips, and avoid aggressive face restoration on already-clean footage.
Do I need to train a model?
Only when volume demands it. For a handful of shots, prompt-level reference fusion plus approved keyframes is faster to set up and easier to revise. Train an adapter when a character recurs across many episodes and needs to hold up under close-ups.
How do I fix one shot that drifted?
Do not fix it inside a finished sequence. Regenerate the keyframe with the correct reference pack, confirm it against the hero frame, then animate again. Patching a drifted clip with editing tricks usually leaves an artefact that reads worse than the drift.
Can I clone a real person's likeness?
Only with explicit permission from that person, and with awareness of the rules in your distribution market. Written consent, clear scope of use, and a review step before publishing are the minimum standards for anything resembling a real, identifiable individual.
Do 3D renders help as references?
Yes, considerably. A simple posed 3D or stylised model gives you perfect multiview coverage under identical lighting, which is exactly what fusion wants. Use renders to supplement photography-style references, then apply the scene grade afterward so the capture layer does not get baked into the identity.
The takeaway
Consistent character cloning is mostly discipline. Curate a small, coherent reference pack. Separate identity from lighting and wardrobe in your mind and in your files. Approve stills before you animate, review cuts rather than individual shots, and post-process everything through one pipeline. Models will keep improving at holding a face steady, but the teams that get reliable results will be the ones treating identity as a managed asset — versioned, documented, and checked at every gate.


