Consistent characters are the difference between a demo clip and a story people can follow. When your protagonist's jawline, jacket, and hairline quietly shift between shots, viewers stop tracking the plot and start spotting the glitch. Multi-image reference fusion — feeding several curated images of the same character into a generative video pipeline — is currently the most practical way to keep identity stable across scenes.
This guide is written for creators who already know how to generate a single good clip and now need a whole sequence that holds together. It covers the mechanics, the reference kit, a repeatable production workflow, prompt structure, troubleshooting, and a continuity checklist you can reuse on every project.
Why AI Video Loses Characters Between Shots
Generative video models are trained to produce plausible motion and plausible imagery. They are not trained to remember that your hero has a scar above the left eyebrow unless something in the pipeline explicitly carries that information forward. Every new generation is a fresh sample from a probability distribution, and small sampling differences compound across shots.
The three drift patterns you will see first
Identity drift. The face stays roughly the same type but the specific person changes — eyes slightly wider apart, nose longer, cheekbones higher. This is the most damaging form because it is subtle enough to survive a casual review and obvious enough to break a sequence.
Wardrobe drift. Colors desaturate or shift hue, fabric texture changes from denim to synthetic, and small details like buttons, stitching, or a collar shape disappear. Wardrobe drift is usually a symptom of insufficient text description combined with weak references.
Prop and set drift. A mug changes size, a car changes model, a room's window moves. Less common in single-shot generation, extremely common in sequences assembled from separate prompts.
Why a text prompt can describe but not lock a face
Text tells a model what category of person to draw. A prompt like "a woman in her thirties with dark curly hair wearing a green field jacket" narrows the space enormously, but thousands of faces satisfy that description. Text cannot encode the exact geometry of a specific face, the precise spacing of features, or the exact shade of a fabric under a specific light. Images can. That is the entire premise behind reference-based generation.
How Multi-Image Reference Fusion Actually Works
Multi-image fusion means conditioning a generation on more than one reference image at once, rather than on a single portrait or a text description alone. The model receives several views of the same subject and is asked to synthesize something new that is consistent with all of them.
Single reference versus multi-reference generation
A single reference is fast and works well when the shot is close to the reference: same angle, same lighting, same distance. The moment the camera moves — profile, three-quarter turn, low angle, back of the head — a single reference leaves too much unspecified and the model improvises. Multi-reference conditioning gives the model a small 3D intuition: it has seen the face from several angles and can interpolate instead of inventing.
The pipeline in plain language
Most modern pipelines follow a similar shape:
- Encode each reference image into an embedding that captures identity, color, and texture.
- Weight and blend those embeddings, sometimes with per-image weights so your best portrait counts more than a blurry profile.
- Cross-attend the blended identity signal against the denoising video latent, so every frame is nudged toward the same identity.
- Propagate temporally so frame 2 and frame 90 are pulled toward the same identity anchor, not two independently sampled faces.
- Decode into frames, then upscale or interpolate to your target resolution and frame rate.
The practical consequence: the more consistent your references are in identity and the more varied they are in angle, the stronger the identity anchor becomes.
How many references is enough
Three is the floor. Five to eight is the productive range for a hero character. Beyond ten, returns flatten and you start importing contradictory information — different lighting, different ages, different styling — which the model averages into a blurry, generic face. Quality and diversity beat quantity every time.
Assembling a Character Reference Kit
Build the kit once, reuse it for the entire project. Treat it like a cast member's contract: fixed, versioned, and not casually edited mid-production.
The minimum viable set
- One straight-on portrait, neutral expression, even light
- One three-quarter view, slight smile
- One profile, both if the character turns often
- One full-body shot establishing proportions and posture
- One detail shot of a signature feature: scar, tattoo, jewelry, glasses
- One shot in the character's primary wardrobe
- One shot in low or dramatic light if the story requires it
Lighting and angle discipline
Keep the lighting in your reference set reasonably neutral. If half your references are lit by warm candlelight and half by cold overcast daylight, the model learns that skin tone is ambiguous. Save dramatic lighting for the prompt, not the reference kit.
Angles should be genuinely different, not cosmetically different. Five near-identical selfies give you the same information five times.
Reference images that sabotage the model
- Heavy filters or beauty retouching that flatten skin texture
- Motion blur or compression artifacts from screenshots of video
- Extreme expressions that distort bone structure — a wide laugh changes the whole face
- Occlusions: hands on faces, hair across the eyes, scarves over chins
- Multiple people in a frame, which can cause the model to blend two identities
- Conflicting ages — one reference from a decade ago and one from today
Spend the extra twenty minutes cropping, color-correcting, and de-duplicating. It saves hours of regeneration later.
A Repeatable Scene-by-Scene Workflow
This is the core of the guide. Follow it in order and character consistency stops being luck.
Step 1: Script, beat sheet, and shot list
Before generating anything, write the shot list with continuity flags. For each shot, note the character's wardrobe, injuries, props, time of day, and emotional state. Continuity errors are cheaper to fix on paper than in generation.
Step 2: Lock the character sheet before production
Generate and approve a character sheet — a grid of the character in the reference angles and the primary wardrobe. Freeze it. Any later change invalidates every clip you have already produced.
Step 3: Generate keyframes, not clips, first
Generate still keyframes for every shot before animating anything. Still images are cheap and fast, and they let you review the whole sequence for identity consistency at a glance. Line up all your keyframes in a grid; the drift becomes obvious immediately.
Step 4: Animate from approved keyframes
Now take each approved still into image-to-video generation with the same multi-image reference set attached. Because the keyframe already locks identity, the video model only has to solve motion, not identity.
Step 5: Run a continuity QA pass
Watch the assembled sequence at normal speed once, then frame-by-frame through every cut. Mark each issue as identity, wardrobe, prop, or lighting. Fix at the lowest level possible — often a single regenerated shot, not a rebuilt scene.
Prompt Structure That Supports Fusion
References do most of the work, but prompts control what the model does with them. Use a five-block structure so you can vary one thing at a time.
The five-block prompt
- Subject block — who, referencing the character sheet explicitly ("the woman from the reference images, same face, same hair")
- Wardrobe block — exact garments, colors, materials, wear state
- Environment block — location, time of day, weather, background elements
- Camera block — shot size, lens feel, angle, movement
- Light block — key direction, quality, color temperature, contrast
Keeping these blocks stable across shots and changing only what the story requires is what produces a sequence that feels shot rather than sampled.
Negative prompts and drift vocabulary
Negative prompts are often underused. Useful entries include: different face, changing features, inconsistent hair, wardrobe change, extra fingers, morphing, flicker, warping background, and text artifacts. Don't overload the negative list — ten focused terms beat forty random ones.
Managing Wardrobe, Age, and Prop Changes
Stories require change. The trick is making change intentional and trackable.
Wardrobe swaps without identity loss
Keep the face references constant and describe the new outfit only in text, or supply a separate wardrobe-only reference. Never swap the portrait references when only the costume changes — that is how identity gets reset.
Aging, injury, and transformation arcs
Create stage-specific character sheets: "act one," "act two," "finale." Each stage gets its own locked reference set that inherits the base identity and adds the change — a scar, greyer hair, torn clothing. Think of it as a costume continuity bible for a film.
Recurring props and sets
Props deserve the same treatment as faces. Give a recurring object its own two or three reference images and reattach them every time it appears. A consistent pocket watch or vehicle is a surprisingly strong continuity cue for viewers.
Troubleshooting: Symptom, Cause, Fix
| Symptom | Likely cause | Fix |
|---|---|---|
| Face changes at cuts | Each shot generated in isolation | Reattach the same multi-image reference set to every shot |
| Face looks generic | References too similar or too heavily filtered | Add profile and three-quarter views, remove retouched images |
| Wardrobe color shifts | Color described only in text | Add a wardrobe reference or use hex-level color language |
| Flicker within a clip | Too many conflicting references | Reduce to five strong references, lower weights on weak ones |
| Background morphs | Prompt describes background loosely | Lock background with a separate scene reference |
| Hands warp | Small in frame, low detail | Frame hands larger or hide them; regenerate at higher resolution |
| Character looks older/younger | Mixed-age references | Rebuild the kit from one photo session or one consistent era |
Most consistency problems trace back to reference hygiene, not model limitations. Audit your references before you blame the tool.
Choosing Tools and Setting Decision Criteria
Tools change quickly, so choose against criteria rather than feature lists.
What to evaluate
- Reference capacity — how many images can you attach, and can you weight them?
- Temporal stability — does the model hold identity across a long clip or drift at second six?
- Motion realism — hands, faces, and walking cycles are the usual weak points
- Control surface — can you supply a keyframe plus references plus a camera instruction?
- Iteration speed — a fast, decent model you can regenerate ten times beats a slow, brilliant one
- Resolution and length — match the model to your target delivery format
- Character lock features — dedicated character or subject-lock functions reduce prompt work
- Post-production friendliness — clean output that composites well in an editing suite
A hybrid stack that works
A reliable setup looks like this: generate and lock character sheets in a strong text-to-image model, produce keyframes there too, then animate with an image-to-video model that accepts multiple reference images. Finish in a conventional editor for grading, sound, and cut rhythm. Tools worth evaluating in that pipeline include Runway, Kling, Luma, Pika, Sora, Midjourney, FLUX, Stable Diffusion with ComfyUI, plus DaVinci Resolve or After Effects for finishing. Test two or three on your own character sheet before committing — the model that handles your specific face is the right one, not the one with the best marketing.
Pre-Export Continuity Checklist
Run this before every export. It takes fifteen minutes and prevents most reshoots.
- Character sheet version matches every clip in the sequence
- Wardrobe state matches the script's continuity notes at each cut
- Injuries, props, and accessories appear and disappear on schedule
- Hair length and style consistent within each stage
- Skin tone and lighting temperature consistent across adjacent shots
- Continuous actions match frame-to-frame where shots are joined
- Backgrounds consistent when the camera returns to a location
- Color grade applied after consistency is verified, not before
- Audio and dialogue timing checked against lip movement
- Full-sequence watch at normal speed, then one frame-by-frame pass
FAQ
How many reference images should I use for a character?
Five to eight well-chosen images: a straight portrait, a three-quarter view, a profile, a full body, a detail shot, and one in primary wardrobe. Add more only if the new images add genuinely new information.
Can I keep a character consistent across different art styles?
Yes, but build separate reference kits per style. A photoreal character sheet will push a stylized render toward realism. For animation or illustration, use a stylized reference set with the same identity logic.
Why does my character look fine in stills but drift in video?
Still generation samples once; video samples across many frames and small errors accumulate. Fix it by generating a locked keyframe first, then animating from that keyframe with the same references attached.
Do I need a different workflow for two characters in one shot?
Yes. Generate each character separately, composite them into a keyframe, then animate. Allowing a model to invent both identities simultaneously is the fastest route to identity blending.
How do I handle a character who changes costume often?
Split identity references from wardrobe references. Keep three to five face images constant and supply wardrobe as a separate reference or a precise text block. Never swap the face kit to change clothes.
Is it worth regenerating a shot for a small continuity error?
If the error appears within the first two seconds or involves the face, yes. If it is a background detail in a fast cut, a grade or a slight crop can often hide it more cheaply.
What is the single biggest cause of inconsistency?
Inconsistent references. Mixed lighting, mixed ages, mixed filters, and duplicate angles confuse the identity signal. A clean, deliberate character kit solves more problems than any prompt trick.
How do I keep consistency across a long series rather than one video?
Version your character sheets and store them with the project files. Treat each stage of the story — and each season — as a sheet version, and record which clips were generated against which version so you can rebuild anything later.
The short version: lock your references, generate keyframes before clips, keep prompts structured, and audit continuity before you export. Multi-image fusion will not fix a chaotic workflow, but it will absolutely reward a disciplined one.

