Why AI video characters drift between shots
Ask any creator who has finished a narrative AI video what the hardest part was, and the answer is rarely the visuals. It is keeping the same face on screen for twelve shots in a row.
A text-to-video model does not know who your character is. It knows what your prompt sounds like relative to the millions of clips it was trained on. When you describe "a woman in her thirties with auburn hair and a grey coat," the model samples a plausible face, a plausible coat, and a plausible lighting condition. Run the same prompt again with a different seed, a different camera angle, or a slightly reworded sentence, and you get a different plausible person. Multiply that across every shot in a sequence and the audience feels something is wrong even if they cannot name it.
This is the consistency problem, and it is the single biggest bottleneck between "impressive demo" and "watchable film." It shows up at three distinct layers:
- Identity drift. Facial structure, eye spacing, jawline, hairline, and skin tone shift between shots. Sometimes it is subtle, sometimes the character becomes a different person entirely.
- Style drift. The rendering style, film grain, color grade, and lens character change from shot to shot, so the sequence feels assembled rather than directed.
- Motion drift. The character holds together in a still frame but deforms once they walk, turn, or speak — ears melt, coats change color mid-pan, hands multiply.
Multi-image fusion is the family of techniques designed to solve the first layer, and when used well, it stabilizes the other two as a side effect. The idea is simple to state and demanding to execute: instead of describing a character with words, you describe them with several images, and you let the model blend those images into a single reusable identity.
What multi-image fusion actually does
Multi-image fusion means conditioning a generative model on more than one reference image of the same subject, then merging the extracted information into one coherent identity representation. It is not one algorithm. It is a set of strategies that different tools implement in different places in the pipeline.
Feature encoding: turning a face into a vector
The first step is extraction. A vision encoder looks at each reference image and produces embeddings that capture what makes that face recognizable: geometry of the face, distance between features, eye and hair color, skin undertone, and to some degree texture and grooming. The encoder deliberately discards what is irrelevant — background, camera angle, JPEG noise — and keeps what is stable across images.
Good encoders separate identity from everything else. A face photographed in harsh sunlight and the same face photographed under soft studio light should produce nearly identical identity embeddings even though the pixels are wildly different. When your tool cannot do that separation, you get a model that faithfully reproduces the lighting of your reference photo in every shot, which looks like a continuity error.
Fusion strategies: early, late, and attention-based
Once you have multiple embeddings, you have to combine them. Three approaches dominate:
Averaging (late fusion). The simplest approach: encode each reference, average the vectors, condition on the result. Fast and cheap, but the average of five photos is often nobody in particular. Averaging smooths distinctive traits — a slightly crooked nose and a strong jawline both regress toward the mean. Use averaging when your references are near-identical and you mainly want noise reduction.
Concatenation (early fusion). All references are passed into the model together, usually with cross-attention so the model can consult different images for different parts of the generation. This is more expressive: the model can take hair shape from one image and jawline from another. It is also more fragile, because the model can weight the wrong reference for the wrong region. Concatenation rewards carefully curated reference sets.
Weighted or learned fusion. The tool assigns or learns a weight per reference, so a clean front-facing portrait dominates while a profile shot contributes less. This is usually the best-behaved option in production, because it lets you include imperfect references without poisoning the identity.
Where fusion sits in the pipeline
Fusion can happen at three points, and knowing which one your tool uses tells you a lot about its behavior:
- Prompt-level. The references are converted into a text-like descriptor that gets merged into your prompt. Weakest identity retention, but works with almost any model and rarely breaks composition.
- Adapter-level. A small adapter network injects identity into a frozen base model. Strong identity retention, controllable strength, and typically the friendliest to work with because you can dial influence up or down per shot.
- Training-level. The identity is baked into a small fine-tuned model or embedding. Strongest retention and most consistent style, at the cost of setup time and one model per character.
Most working creators use a combination: adapter-level conditioning for speed and control during exploration, then a trained identity for the hero character in the final sequence.
Building a reference set that survives every shot
Fusion quality is capped by reference quality. A perfect fusion algorithm fed five inconsistent selfies will produce a confidently inconsistent character. Budget real time for this step; it pays back more than any prompt trick.
The minimum viable reference set
For a character who appears in more than a handful of shots, aim for eight to fifteen images:
- One clean front-facing portrait, neutral expression, even lighting
- One three-quarter view, both left and right if possible
- One profile view
- One shot with the mouth open or mid-speech
- Two to three shots at different emotional registers (neutral, smiling, serious)
- One full-body or three-quarter-body shot to lock proportions and wardrobe
- One shot under warm or dim light so the model learns the face is not married to one color temperature
If you only have three images to work with, use front, three-quarter, and a mild expression variation. Never use three near-identical portraits; you will get a rigid identity that shatters the moment the camera moves off-center.
Anglee: keep technical consistency across references
Angles matter, but so does the technical envelope. Mixing a phone photo with a 200mm studio portrait and a screenshot from a different project gives the encoder conflicting signals about skin texture and lens distortion. Where you can, keep the following similar across the set:
- Resolution and sharpness
- Color temperature and white balance
- Focal length feel (avoid extreme wide-angle selfies)
- Absence of heavy filters or beauty-mode processing
- Neutral background where possible
What to leave out
Some references actively hurt. Remove any image with:
- Sunglasses, masks, or hands covering the face
- Heavy stylistic makeup that changes facial structure
- Motion blur or occlusion
- Two people where one could be mistakenly extracted
- Aggressive HDR or heavy skin smoothing
- Watermarks and text overlays
If your character wears glasses, that is a design decision — include them consistently in at least most references, or exclude them from all and add them in post. Mixing "sometimes glasses, sometimes not" teaches the model that glasses are optional noise, and you will get flickering frames.
A practical fusion workflow from script to final cut
The following workflow assumes image-first generation: you produce keyframes with a strong image model, then animate them with a video model that accepts identity conditioning. It is the most reliable path today because it gives you a checkpoint to review before you spend time on motion.
Step 1 — Write the shot list before you generate anything
This sounds like film school advice, but it is a technical requirement. Your shot list tells you which angles you need consistent, which is the list of angles your reference set must cover. If shot seven is a low-angle close-up and none of your references include a low angle, that shot is where the character will break.
For each shot, note: framing, camera height, lens feel, lighting direction, wardrobe state, and whether the character speaks. Keep it in a spreadsheet.
Step 2 — Package references per character
Create one folder per character. Inside, keep a refs/ subfolder with the numbered reference images and a short text file describing the character in words: age, build, hair, wardrobe, distinguishing marks. You will paste that description into every prompt. This is your identity anchor.
Include a bad-refs/ folder for images you decided against. Six weeks later you will wonder why you excluded one, and having it documented saves hours.
Step 3 — Generate a keyframe sheet
Generate a still for every shot before animating any of them. This is the single highest-leverage habit in AI video production. Lay the stills side by side, scaled to the same height. You are looking for:
- Facial proportions holding across angles
- Wardrobe and hair state matching between adjacent shots
- Consistent color grade and contrast
Fix problems here. Regenerating a still costs seconds; regenerating a five-second animated shot costs minutes and sometimes a queue slot.
Step 4 — Animate with identity lock engaged
Once the keyframe sheet passes review, animate each still. Keep the identity conditioning strength at a moderate setting. Too low and the video model winds up drifting toward its own prior; too high and the character becomes stiff, with minimal expression range and a plastic quality.
If your tool exposes separate controls for identity and motion, treat them as independent dials. Identity stays constant across the whole sequence; motion can vary per shot.
Step 5 — Run a continuity pass
Watch the assembled sequence at normal speed, then at half speed, then frame-by-frame on the cuts. Most consistency failures are invisible at normal speed on your own monitor but obvious on a phone screen at half speed. Keep a fix list with timestamps and be ruthless about re-rendering the two or three worst shots rather than accepting them.
Control techniques: prompts, seeds, and adapters
Fusion handles identity, but the surrounding controls determine whether it holds. These are the levers that matter most.
Identity anchors in prompts
Rewrite your character description as a fixed string and paste it verbatim into every prompt. Do not paraphrase between shots. Small wording changes — "grey wool coat" versus "long grey coat" — nudge the model in different directions, and when identity conditioning is already under strain, that nudge is enough to shift the wardrobe.
Structure the prompt so the anchor comes first, then the shot-specific details:
[IDENTITY ANCHOR: mid-thirties woman, auburn shoulder-length hair,
freckles across nose and cheeks, grey wool coat, dark green scarf]
Shot: medium close-up, low angle, evening street light from the left,
slight rain, shallow depth of field
Seed discipline
Where the tool allows it, lock the seed per character rather than per shot. Some pipelines respond well to a single seed shared across the sequence, which reduces stylistic jitter. Others respond better to per-shot seeds with strong identity conditioning. Test both on a three-shot sample before committing to a full sequence.
Motion versus identity tradeoff
Every video model has a limited budget for how much information it can hold constant. Push motion hard — fast turns, running, complex hand gestures — and identity fidelity drops. This is not a bug you can prompt your way out of. Design shots around it: if a character must sprint through frame, keep them mid-distance and in motion blur, where facial detail is not being judged.
Camera moves and framing changes
Dolly-ins and slow arcs are friendly to identity. Whip pans and rapid cuts between wildly different focal lengths are not. When a shot demands a big change in framing, split it into two generations and cut between them rather than asking the model to travel the whole distance in one take.
Negative guidance
Use negative guidance for structural errors, not for identity. Phrases like "extra fingers," "warped face," and "duplicate limbs" are useful. Negative guidance that tries to describe a face you do not want rarely works and often degrades the result.
Troubleshooting common consistency failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Face changes shape between angles | Reference set too narrow | Add profile and three-quarter references |
| Character looks like a sibling, not the same person | Averaging fusion flattening distinctive traits | Switch to weighted fusion or a trained identity |
| Wardrobe color shifts | Prompt wording inconsistency | Lock the identity anchor string verbatim |
| Background lighting bleeds onto face | References shot under strong colored light | Add neutrally lit references |
| Character goes plastic when dialed up | Identity strength too high | Reduce strength, add expression range in references |
| Style changes at every cut | No shared seed or style reference | Lock seed per character, add a style reference frame |
| Hands and limbs deform in motion | Too much motion in one generation | Shorten the shot or increase distance from camera |
| Flicker across frames | Mixed identity weights between adjacent shots | Standardize settings across the sequence |
Choosing a tool stack: decision criteria
Tool marketing rarely answers the questions that matter for consistency work. Evaluate stacks against these criteria instead.
Reference count and quality handling. How many images can you supply, and does the tool weight them intelligently or naively average them? Tools that let you include a weak reference without damaging the result are worth a lot.
Strength control granularity. Can you dial identity influence per shot, or is it on or off? Per-shot control is essential for sequences with varied framing.
Keyframe-first support. Does the pipeline let you generate and review stills before committing to motion? If not, you will burn enormous time on rejected animation.
Angle generalization. Test a reference set of only front-facing portraits and then ask for a profile shot. Tools that extrapolate well save you from hunting for perfect references.
Seed and parameter exposure. If you cannot reproduce a good result, you cannot build on it.
Batch and sequence tooling. For anything longer than a minute, you need to generate dozens of shots with the same settings. Manual repetition introduces errors.
Export and handoff. Resolution, frame rate, alpha channel support, and whether you can export a still sequence for finishing work.
Reliability. A tool that produces a great result one time in five is worse than a modest tool that produces an acceptable result every time. Run the same prompt five times and watch the variance.
A common practical setup mixes categories: an image model with multi-reference conditioning for keyframes, a video model with identity conditioning for animation, and a lightweight upscaling and grading step at the end. Trying to do everything in one model is convenient but usually means compromising on whichever layer you cannot control.
Scaling consistency across a series and a team
Solo creators can hold a character in their head. Teams cannot. If more than one person generates shots, you need infrastructure.
Asset library and naming conventions
Adopt a naming scheme that encodes character, shot, and version: char-aria/shot-014/keyframe_v03.png. Keep approved assets in a read-only folder. When someone needs to fix a shot, they work in a scratch folder and only promote the result after review.
Prompt templates
Store the identity anchor and the standard negative guidance in a shared template file. Every generated shot starts from the template. This eliminates the most common source of drift in team projects: two people describing the same character two slightly different ways.
Review gates
Three gates work well: keyframe approval, animation approval, and final assembly approval. Each gate has a short checklist derived from the troubleshooting table above. Keep the checklists visible — consistency review is repetitive, and repetition makes people skip steps.
Version control for identity
When you improve a character's reference set or train a new identity model, version it. Note which shots were generated with which identity version. Otherwise a mid-project improvement creates a visible seam where the character suddenly gets better-looking.
FAQ
How many reference images do I really need?
Five is the practical floor for a character in a short sequence; eight to fifteen is comfortable. Beyond twenty, returns flatten and you risk introducing contradictory signals unless the images are very consistent.
Can I get consistent characters from a single image?
Sometimes, for short sequences with limited angle variation. Expect identity to weaken whenever framing changes dramatically. A single reference is a starting point, not a foundation.
Should I train a dedicated identity model for each character?
If the character appears in a multi-episode series or a long film, yes. Training costs setup time but pays back in stability, and it lets you rebuild the character months later without digging through old references.
Why does my character look fine in stills but drift in video?
Video models devote capacity to motion, leaving less for identity. Shorter shots, slower movement, and moderate identity strength usually fix it. Also check whether the video model is receiving the same references as the image model.
Does wardrobe count as identity?
It does when the audience notices. Treat costume changes as deliberate story beats, not as accidents. Lock wardrobe in the identity anchor for every shot where the outfit should not change.
What is the fastest way to test if a tool handles fusion well?
Give it three references of the same person from different angles and ask for a profile view. If the result is recognizably the same person and the background does not carry over from the references, the fusion is working.
How do I handle a character who changes appearance on purpose?
Split them into separate identity versions — before and after — and generate each block of shots with the correct version. Avoid asking the model to interpolate the transformation unless that transformation is itself the shot you are building.
Key takeaways and a working checklist
Multi-image fusion is not a magic switch. It is a set of encoding and blending techniques whose output is only as good as the references you feed them and the controls you apply afterward. The creators who get reliable results are not using secret tools; they are running a disciplined process.
Before your next sequence, confirm each of these:
- You have a shot list with angles, framing, and lighting documented per shot.
- Each character has eight or more references covering front, three-quarter, and profile views, with expression variety and at least one neutral-light image.
- Excluded references are archived with a note explaining why.
- The identity anchor text is stored in a template and pasted verbatim into every prompt.
- A keyframe sheet exists and passed review before any animation was generated.
- Identity and motion controls are set deliberately, not left at defaults.
- The assembled sequence has been reviewed at half speed with a timestamped fix list.
- Approved assets live in a versioned, read-only folder with a naming convention your team understands.
Work through that list once and the difference is immediate: fewer re-renders, fewer reshoots, and a character the audience can actually follow from the first frame to the last.

