A viewer will forgive a lot. They will forgive a slightly soft background, a camera move that does not quite land, even a soundtrack that is a little generic. What they will not forgive is a character whose face changes between shots. The moment the jawline shifts, the eyes change color, or a jacket becomes a different jacket, the audience stops watching a story and starts watching a tool.
That single problem — visual identity that survives across scenes — is what separates a clip that demos well on social media from a sequence that can carry a narrative, an ad campaign, or an episodic series. The good news is that modern multi-image reference and feature-merging workflows have made consistency an engineering problem rather than a lottery. This guide walks through the full pipeline: preparing references, writing locked identity prompts, generating a keyframe set, animating it, and quality-checking the result.
Why Character Consistency Breaks Down in AI Video
Most generative video models are trained to make each frame look plausible in isolation. Plausibility is not identity. When you prompt "a woman in a red coat walks through a rain-soaked street," the model samples from an enormous space of women in red coats. Take that prompt to a second scene and it samples again — landing on a different woman. Nothing is broken; the model simply has no reason to remember the first one.
In practice, three separate kinds of drift show up, and they need different fixes.
Identity drift
The face, hair, eye color, skin tone, and body proportions change. This is the most visible failure and usually the first one you notice. It gets worse when the camera angle changes: a model that can hold a frontal face may completely reinvent the person in profile, because profile faces are a smaller slice of the training distribution.
Style and grade drift
Lighting, color temperature, lens character, and film grain wander from shot to shot. Even if the face holds, a warm interior followed by a cool, contrasty exterior reads as two different productions. Continuity of look is nearly as important as continuity of face, and it is easier to fix in post — but it is much cheaper to fix at generation time.
Continuity drift
This covers wardrobe state, props, and story logic. The coat is buttoned in one shot and open in the next. The coffee cup is full, then empty, then full again. A bruise appears two scenes before the fight. Continuity is a writing and shot-planning problem as much as a generation problem, and no reference image will save you from a scene list that contradicts itself.
How Multi-Image Reference Actually Works
A single reference image gives a model one view of a person. Multi-image reference gives it several, and — critically — a way to fuse them into a consolidated identity representation rather than treating them as separate, competing inputs.
Mechanically, each reference image is passed through an image encoder that converts it into a set of feature embeddings. Those embeddings are then injected into the generation process, usually through cross-attention layers, so that the model's sampling is conditioned not just on your text but on the visual characteristics of your references. When several references are supplied, the system blends their features — a kind of feature-level averaging of the face cluster — so that shared traits (bone structure, hairline, eye spacing) reinforce each other while view-specific traits (the exact shadow under the chin in one photo) get diluted.
That blending behavior is why reference selection matters so much. If you feed the model five photos taken in identical lighting from identical angles, you get a very strong but very narrow identity. If you feed it one clean front, one three-quarter, one profile, one full body, and one costume shot, you get a broader, more robust identity that holds up when the camera moves.
What the reference slots are actually good for
- Front-facing, evenly lit portrait: anchors facial geometry and eye color.
- Three-quarter view: teaches the model how the face behaves when it rotates.
- Profile: prevents the common "invented nose" problem in side shots.
- Full body: locks height, build, and limb proportions — important for wide shots.
- Costume or wardrobe flat-lay: locks garment color, cut, and trim details.
- Expression set: helps the model produce emotion without rebuilding the face.
The role of the merge step
The merge step is where you decide how much authority each reference carries. When you merge references, you are effectively asking the model to treat them as one person rather than as a mood board. If your tool exposes weights, give the frontal portrait the highest weight, the three-quarter second, and background or prop references the lowest. If it does not expose weights, then curate by subtraction: include only the references you are willing to have averaged into the result.
Step 1: Build a Character Bible Before You Generate Anything
Professional animation studios do not start with a shot; they start with a model sheet. Do the same, even for a ten-second clip. Create a folder per character with eight to twelve images and a short text file that describes them in words.
A workable character bible contains:
- A neutral portrait, frontal, soft even light, no heavy shadows across the face.
- Two three-quarter views, one from each side.
- A left and right profile.
- One full-body shot in the default wardrobe, at a consistent focal length.
- Three to five wardrobe variations in identical lighting.
- Two expression extremes — a broad smile and a neutral-to-serious look.
- Two lighting variants — warm tungsten and cool daylight — so the model learns what is identity and what is light.
Keep backgrounds plain and consistent. A busy background is a signal the model may absorb, which is exactly what you do not want when the actual scene is a crowded train platform. Crop tightly enough that the face occupies a large portion of the frame, and keep source images sharp — upscale a small reference rather than feeding a blurry one, because blur propagates into the generated face.
Naming and storage discipline
Use a strict naming convention: character_scene-shot_ref-role.png, for example mara_sc01-01_ref-front.png. Store the reference images, the merged identity preset, the exact prompt text, the seed value, and the model name alongside each other. Six scenes later, when a producer asks for a reshoot, you will be able to reproduce the shot in minutes instead of guessing.
Step 2: Write a Locked Identity Descriptor
The second half of consistency is textual. Even with strong image references, your prompt is a set of instructions the model will follow — and vague instructions give it room to improvise.
Split every prompt into four blocks, in this order:
- Identity block: age, ethnicity, build, hair color and texture, eye color, distinguishing features, and a short sentence describing the character's presence.
- Wardrobe block: garment names, colors, materials, and state (buttoned, sleeves pushed up, tie loosened).
- Scene block: location, time of day, weather, background elements.
- Camera block: shot size, lens feel, angle, movement, and depth of field.
Here is the important rule: the identity block must be byte-for-byte identical across every shot of a scene, and ideally across the whole project. Do not paraphrase it. Do not improve it halfway through. If you find yourself wanting to change the identity block, that is a signal to restart the shot list rather than to edit mid-stream.
A negative prompt that earns its place
Negative prompts should target the failure modes you actually see, not an encyclopedic list. A practical starting set for character work includes: face morphing, distorted eyes, asymmetric pupils, extra fingers, changed hairstyle, different person, plastic skin, over-smoothed texture, and inconsistent wardrobe color. Add project-specific items — "no glasses," "no hat" — because accidental accessories are a common and maddening form of drift.
Step 3: Generate a Keyframe Set, Not a Video
The single most effective workflow decision is to separate still-image generation from motion generation.
First, generate one locked keyframe per shot using your multi-image reference set. Iterate on these stills until every one of them reads as the same person in the same world. This stage is fast, cheap, and easy to judge — your eye catches identity drift in a still far more reliably than in motion, because motion masks small distortions.
Second, animate each approved keyframe with image-to-video, so the model's job becomes "move this person" rather than "invent this person." The reference image carries identity; the motion prompt carries performance. This division of labor is what makes the pipeline reproducible.
Building the shot list
Write your shot list as a table with one row per shot and columns for scene number, shot size, description, wardrobe state, time of day, and duration. Most models produce their best work in short increments, so plan on clips of roughly four to eight seconds and build longer sequences through editing. Mark every shot that shares a location so you can batch them and reuse a consistent lighting prompt.
Step 4: Animate Each Shot From Its Locked Frame
Motion prompts should describe movement, not appearance. If you re-describe the character's face here, you invite the model to redraw it. Keep motion prompts to action, camera behavior, and atmosphere: "she turns toward the window, camera slowly pushes in, rain streaks the glass, subtle handheld sway."
Practical rules that reduce drift during animation:
- Keep the first frame identical to your approved keyframe; never let the tool re-generate frame one.
- Prefer one dominant action per clip. Two actions in six seconds produces mush.
- Keep camera moves simple. Dolly, pan, and slight handheld are reliable; whip pans and complex arcs are not.
- Use consistent motion vocabulary across a scene so cuts feel intentional.
- Reserve a longer, slower shot for any moment where the audience must read emotion on the face.
Cutting for continuity
In editing, cut on motion wherever possible. A cut during a hand gesture or a turn hides small continuity differences far better than a cut on a static frame. Where you must cut between static shots, insert a reaction shot, a prop insert, or an establishing wide — these are the standard tools editors have used for decades, and they work just as well on generated footage.
Step 5: Handling Deliberate Variation
Consistency does not mean a character can never change. It means changes are chosen, not accidental. Adopt a one-variable rule: each shot changes exactly one attribute — wardrobe, lighting, or location — while everything else stays locked.
- Wardrobe changes: generate a new reference image of the character in the new outfit using the same merged identity preset, then add it to the bible. Do not simply edit the wardrobe block and hope.
- Time-of-day changes: keep the identity block fixed and change the scene and lighting language only. If skin tones shift too warm, correct in color grading rather than regenerating.
- Age or injury progression: treat these as separate identity presets derived from the original. Create a "character, ten years later" preset and use it consistently from the first shot in which it applies.
- Emotional states: rely on expression references plus motion prompts, not on rewriting facial description. "Jaw tight, eyes wet" is performance; "narrower face, thinner lips" is a new character.
Choosing Your Tooling: Decision Criteria
Model choice matters less than pipeline discipline, but the differences are real. Evaluate candidates against these criteria rather than against demo reels.
- Reference capacity: how many images can be supplied at once, and do weights exist?
- Identity fidelity versus motion realism: some models hold faces beautifully but move stiffly; others move convincingly and drift. Pick per project type — dialogue-driven scenes favor fidelity, action favors motion.
- Resolution and aspect-ratio support: vertical for social, widescreen for narrative.
- Seed determinism: can you reproduce a shot exactly?
- Iteration cost: time per render, batch size, and whether you can parallelize.
- Character-lock features: dedicated identity presets, trained character adapters, or reference slots inside the video model itself.
- API and automation: if you are producing at volume, scripted generation with fixed parameters beats manual clicking every time.
For most projects, a hybrid strategy works best: use a still-image model with strong multi-image reference for keyframes, and a video model with reliable image-to-video behavior for animation. Training a small character adapter on your bible images is worth the setup time if the character will appear in dozens of shots.
Common Mistakes and How to Debug Them
The face changes every shot. Your identity block is being paraphrased, or your references are too similar to each other. Diversify angles and freeze the text.
The face is right but the person looks like a different build. Add full-body references. Wide shots expose proportions that portraits never teach.
Wardrobe color shifts. Reference the garment directly in an image, and add a negative prompt for the wrong color. Color words alone are weak conditioning.
Everything looks over-smooth and plastic. Reduce the number of high-gloss reference photos, add texture references (fabric close-ups, skin detail) and a negative prompt against beauty-filter finish.
Motion is fine but the character seems to float. Check the keyframe for improbable contact with the ground, then shorten clip duration and simplify the move.
Scene looks like two different films. Create one lighting prompt per location and reuse it verbatim; then apply a single color grade across the whole timeline.
Reproducing a shot is impossible. You did not save the seed, prompt, and reference set. Start saving them — today's convenience is next week's crisis.
A Practical QA Checklist
Before you call a sequence finished, check each item on every shot:
- Same face geometry, eye color, and hairline as the approved preset.
- Wardrobe matches the shot list row exactly, including state.
- Lighting temperature matches adjacent shots in the same location.
- No accidental accessories, jewelry, or logo changes.
- Hands and fingers read correctly in any close shot.
- Background elements are consistent with the scene's geography.
- Cut points land on motion or on a deliberate reaction insert.
- Sound design and room tone are continuous across the cut.
- Final grade is applied uniformly, with skin tones protected.
FAQ
How many reference images do I actually need? Four to six well-chosen images cover most needs: a frontal portrait, a three-quarter, a profile, and a full body. Add wardrobe and expression references as the script demands. More is not automatically better — redundant, near-identical references narrow the identity instead of strengthening it.
Can I fix an inconsistent shot in post? Sometimes. Face-swapping or compositing a correct face onto a drifting shot works for brief moments, but it is slow and rarely convincing in close-up. Regenerating from a locked keyframe is usually faster and cleaner.
Should I train a custom character model? If the character appears in more than about twenty shots, yes. Training on your bible images produces a durable identity that no prompt needs to re-explain, and it lets you work with shorter prompts and looser reference sets.
Why does the character look right in stills and wrong in video? Motion models re-sample identity across frames. The fix is not better prompts but a stronger starting frame plus simpler motion. Shorten the clip, reduce camera movement, and keep one action per shot.
How do I keep a whole series consistent across episodes? Treat the character bible and the prompt blocks as versioned assets. Store them in a project repository, tag stable versions, and never edit an identity block mid-episode. Consistency is a file-management discipline as much as a creative one.
What about voice consistency? The same principle applies: choose one voice profile early, document its parameters, and reuse it in every line. Audiences associate voice with identity just as strongly as face, and a mid-project voice change undoes much of your visual work.
Building consistent characters is not about finding a magic model. It is about treating identity as a fixed asset that flows through a pipeline — references in, locked prompts through, keyframes approved, motion layered on top, and a checklist at the end. Do that, and the audience stops noticing the tool and starts following the story.

