Character drift is the quiet tax on every AI video project. You generate a hero shot that looks exactly right, then the next shot gives you a slightly different jaw, a different nose, a different age. Multiply that by twenty shots and you are no longer making a film, you are running a casting session that never ends.
The good news is that consistency is not a matter of luck or of finding the one magic model. It is a workflow problem. Studios that ship coherent AI video treat character identity as a data asset that is defined once, reused everywhere, and checked at every stage. This guide walks through that workflow end to end: how to prepare a reference kit, how multi-image referencing actually influences generation, how to direct scenes so a face survives camera moves, and how to catch drift before it reaches the timeline.
Why AI Video Characters Drift Between Shots
A text-to-video model does not have a memory of your protagonist. It has a probability distribution over pixels, conditioned on whatever text and images you hand it for this generation. Change the prompt, the seed, the aspect ratio, or the camera move, and you change the conditioning. The model then re-samples a face that fits the new conditions. Sometimes it lands close. Often it lands somewhere adjacent and plausible but wrong.
Identity versus style
It helps to separate two things that get conflated constantly.
- Identity is the invariant set of features that makes a viewer say "that's her": face geometry, skin tone, hairline, eye spacing, body proportions, and any signature marks.
- Style is everything that should change between shots: lighting, grade, lens, wardrobe, environment, pose, and mood.
Most drift complaints are actually identity leakage into style or style leakage into identity. When you ask for a dramatic night scene and the model answers by relighting the face into a different person, identity and lighting have been entangled in the conditioning. Your job is to keep them in separate boxes.
What continuity actually means on screen
Audiences are forgiving in specific ways. They will accept a slightly different cheekbone if the hair silhouette, wardrobe, and colour palette match. They will not accept a different silhouette, a different hairstyle, or a face that changes age between cuts. Prioritise in this order:
- Silhouette and hair shape
- Wardrobe and colour anchors
- Face geometry at close range
- Micro-details like freckles and jewellery
This ranking matters because it tells you where to spend generation attempts and where you can stop worrying.
Build a Character Reference Kit Before You Generate
A reference kit is a small, curated image set that represents the character from multiple angles under neutral conditions. It is the single highest-leverage thing you can build, and it takes an afternoon.
What to collect
Aim for eight to fifteen images. More is not better if the images contradict each other.
- Three head angles: straight on, three-quarter, profile.
- Two body framings: full body and mid-shot, in the same outfit.
- Two expressions: neutral and one emotional extreme, usually a smile.
- One silhouette test: a backlit or shadowed frame where only the outline reads.
- Optional variants: a younger version, an older version, or an alternate costume, kept in separate folders.
Image hygiene
Reference quality beats reference quantity. Before you add an image to the kit:
- Crop tightly on the subject and remove distracting background clutter.
- Normalise lighting where possible. A reference shot in harsh green light will pull green into every generation.
- Upscale anything below roughly 1024 pixels on the short edge, then check for artificial smoothing.
- Remove watermarks, text overlays, and heavy filters.
Naming and versioning
Name files so a collaborator can understand them without asking: hero-lineup-v3-threequarter-neutral.png. Keep one folder per character per production. When you change a reference, bump the version rather than overwriting, because you will sometimes need to roll back after a model update changes how it interprets your images.
Locking Identity With Prompts, Seeds, and Reference Weight
The mechanical part of consistency lives in three controls: text, seed, and image conditioning weight.
Write an identity block, then freeze it
Draft a short paragraph describing only permanent features. No lighting, no mood, no camera direction. Something like: "Adult woman, early thirties, oval face, wide-set dark brown eyes, straight nose, defined jaw, shoulder-length wavy black hair with a centre part, medium build." Save it as a reusable snippet. Every prompt for that character starts with the same block, verbatim. Do not paraphrase it between shots. Small wording changes produce small identity changes, which is exactly the drift you are trying to avoid.
Seeds and latent anchoring
In most diffusion-based pipelines, a seed determines the initial noise. Reusing a seed across shots gives you a related starting point, which nudges facial features toward the same neighbourhood of the model's output space. Seeds are not a guarantee, because once you change the prompt substantially the sampling path diverges anyway. Treat a fixed seed as a stabiliser, not a lock.
Stronger locking comes from anchoring the identity in the latent space using image conditioning. Reference-image adapters, identity encoders, and face-region inpainting all work on the same principle: inject features from your reference images into the denoising process so the model has a concrete target rather than a text description of a target.
How much reference weight is too much
Reference influence is a dial, and both ends fail in predictable ways.
| Weight | Typical result |
|---|---|
| Too low | Face drifts shot to shot; the text prompt dominates |
| Balanced | Identity holds while lighting and pose change freely |
| Too high | Face copies the reference pose and lighting; motion freezes into a photo |
Start in the middle, generate three test frames with different lighting, and adjust. If lighting from the reference bleeds into a night scene, lower the weight or mask the conditioning to the face region only.
Multi-Image Fusion Techniques That Hold Up
Multi-image referencing — feeding several images of the same character at once — is where consistency stops being fragile.
Blend references, do not stack them
When you supply multiple images, you are effectively asking the model to find a shared identity across them. That works best when the images agree on the fundamentals. If one reference has a rounder face than another, the blended output will drift toward an average that matches neither. Curate for agreement first, then blend.
Style transfer versus identity preservation
These are separate operations and should be treated as separate passes.
- Identity pass: lock the face and body using your reference kit, neutral environment, neutral light.
- Style pass: apply grade, lens character, film grain, and atmosphere to the approved frame.
Trying to do both in one generation is the most common cause of "the face changed when I asked for cinematic lighting." Approve the face first, then style it.
Handling deliberate changes
Characters age, get injured, change costume, and get soaked in rain. Handle each change with a derivative reference rather than a new prompt.
- Aging: create an aged reference once, verify it, then use it as the identity anchor for that act.
- Costume: keep the face anchor identical and change only the wardrobe description and wardrobe reference image.
- Damage and wetness: apply as a post-pass or a low-weight style pass so geometry stays intact.
Directing Scenes So the Character Stays On-Model
Consistency is partly a cinematography decision. Certain shots are simply easier for a model to keep coherent, and a well-planned scene will need fewer retakes.
Blocking and camera distance
Extreme close-ups amplify every small facial inconsistency. Wide shots hide them. If your script allows it, build scenes so emotionally important dialogue happens in medium shots and reserve extreme close-ups for one or two beats per sequence. When you do go to a close-up, generate it from an approved medium frame rather than from text alone.
Lighting and colour continuity
Define a three-colour palette per scene: key, fill, and practical accent. Write those colours into the prompt in every shot of that scene. Consistent colour language does two things: it keeps the world coherent, and it stops the model from reaching for a dramatically different lighting interpretation that drags the face with it.
Wardrobe as a continuity anchor
Wardrobe is the cheapest consistency signal you have. A recognisable jacket, scarf, or colour block lets viewers track a character even when the face is imperfect. Lock one wardrobe reference image per act and reuse it, and avoid generating costume variations you do not intend to keep.
A Shot-by-Shot Workflow From Script to Final Cut
Here is a production sequence that keeps drift under control at every stage.
1. Break the script into a shot list with continuity columns
Create a simple table: shot number, description, location, time of day, wardrobe state, and emotion. This becomes the single source of truth for prompts and for review.
2. Generate a hero frame per scene, not per shot
For each scene, spend your attempts on one strong, on-model still that represents the character in that scene's lighting and wardrobe. Approve it deliberately. This frame becomes the visual contract for the scene.
3. Animate outward from approved frames
Use image-to-video with the approved hero frame as the first frame or as a strong reference. Generate the rest of the scene's shots from variations of that frame rather than from scratch. This is the single biggest reduction in drift you will see.
4. Review with a contact sheet, not in isolation
Lay the shots out as thumbnails in sequence and look at the whole strip. Drift that is invisible shot by shot becomes obvious in a strip. Flag any shot where the hair silhouette or jawline reads differently from its neighbours.
5. Fix in the smallest unit possible
Regenerate the flagged shot, not the scene. If the flagged shot is a close-up, fix it by inpainting or by face-region correction rather than by rerolling everything and losing the good motion.
6. Finish with a light grade pass
Apply colour, grain, and lens character across the whole sequence at once. A unified grade hides small residual inconsistencies and makes the sequence read as one continuous piece of filmmaking.
Choosing Tooling: What Actually Matters
Features list look identical across products. These criteria separate tools that help you stay consistent from tools that make it harder.
- Reference image support: how many images can you supply, and can you weight them individually or mask them to a region?
- First-frame and last-frame control: can you animate between two approved stills? This is the most reliable consistency technique available.
- Model flexibility: can you switch models mid-project and still reference the same identity asset, or does each model need its own reference set?
- Iteration cost and speed: consistency is an iterative process. If a single attempt takes fifteen minutes, you will stop iterating and accept drift.
- Asset management: does the tool keep references, seeds, and prompt snippets organised per character, or do you maintain spreadsheets by hand?
- Export and handoff: can you pass clean plates and metadata to an editor without losing track of which reference produced which shot?
A useful test: take one character, one reference kit, and three shots in different lighting. Run the same test in two tools and compare how much prompt editing was required to keep the face stable.
Common Mistakes and How to Fix Them
Overloading the prompt
Long prompts with contradictory style words force the model to compromise. Keep the identity block short and stable, and describe style in a separate, clearly separated clause.
Using inconsistent references
If two references disagree about face shape, the blend will wander. Audit your kit for agreement before blaming the model.
Ignoring first-frame discipline
Text-to-video for every shot is the fastest route to a cast of lookalikes. Generate or approve a still first, then animate.
Chasing a perfect face in a wide shot
At a distance, facial detail contributes almost nothing. Spend your attempts on silhouette, wardrobe, and motion instead.
Forgetting aspect ratio and lens language
Changing from 16:9 to 9:16 mid-project reprocesses the framing and often changes the face. Decide deliverables before you generate.
Editing before the sequence is locked
Cutting and rearranging shots before consistency review hides drift and makes the eventual fix more expensive. Lock picture, then lock continuity, then cut.
Quality Control Checklist Before You Publish
Run this pass on every sequence:
- Build a contact sheet of all shots featuring the character, in order.
- Check silhouette and hair shape first, then wardrobe, then face.
- Compare each close-up against the approved hero frame for that scene.
- Verify colour continuity across scene boundaries, not just within scenes.
- Watch the sequence muted. If the character reads as one person without dialogue, consistency is working.
- Watch it at 2x speed. Fast playback exposes flicker and shape pops that normal speed hides.
- Archive the references, seeds, and prompt snippets alongside the project files.
FAQ
How many reference images do I actually need?
Eight to twelve well-curated images are usually enough for a consistent character. Beyond that, the returns flatten unless the extra images cover a genuinely new angle or a new costume state.
Can I keep a character consistent across different models?
Partially. Identity anchors transfer better than seeds or prompts. Expect to rebuild the reference weighting per model and to re-approve one hero frame as your calibration target.
Why does my character change when I change the aspect ratio?
Reframing changes what the model sees and how it distributes detail across the frame. Generate in your delivery aspect ratio from the start, or letterbox rather than reframe.
Is a fixed seed enough for consistency?
No. Seeds reduce randomness at the start of sampling but do not encode identity. They work best as a secondary stabiliser on top of image conditioning.
How do I handle a character who must age across the story?
Build a small set of age-state references — young, present, older — and treat each as its own identity anchor for its act. Do not try to describe aging in words alone.
What is the fastest way to fix one bad shot?
Regenerate from the approved hero frame of that scene with the same prompt and wardrobe state, changing only the motion or camera description. Inpaint the face if the body and motion are otherwise good.
Do I need a storyboard before generating?
You need a shot list with continuity columns at minimum. Full boards help most when a scene involves complicated blocking or multiple characters interacting.
How do I stop lighting from changing the face?
Separate the passes. Lock identity under neutral light, approve the frame, then apply the scene lighting as a style pass. If you must do it in one generation, reduce reference weight and mask conditioning to the face region.
Consistency is not a feature you switch on. It is a discipline of curating references, freezing identity language, animating from approved frames, and reviewing sequences as strips rather than as individual clips. Do those four things and the casting-session feeling disappears from your timeline.




