Character drift is the quiet tax on every generative video project. You finally generate a hero shot that looks right — the face, the silhouette, the jacket, the way the key light lands on the cheekbone — and the very next clip returns a distant cousin. Do that across twenty shots and the series stops feeling like a story and starts feeling like a casting accident.
The Lego pixel method attacks that problem directly. Instead of trying to lock an entire character inside one long, hopeful prompt, you break the character into small, describable blocks, freeze most of them, and let the model improvise only inside tightly bounded zones. The result is not a magic prompt. It is a repeatable architecture that survives new angles, new lighting, new sessions and new collaborators.
Why Character Drift Is the Quiet Tax on Every AI Video Project
Every shot in a generative pipeline is a fresh sample. Even with an identical prompt, seed noise, attention weighting and a hundred intermediate latent decisions push the result in slightly different directions. Add a change of angle, a change of wardrobe, a change of lighting or motion, and the cumulative difference becomes visible: the jawline softens, the hair parts on the other side, the eyes shift a shade, the nose loses its bridge.
Four drivers cause most of the damage:
- Prompt dilution. A long character description competes with scene, camera and lighting instructions. Identity details get buried under atmosphere words, and the model treats them as suggestions rather than law.
- Reference poverty. One reference image gives the model exactly one viewpoint. Change the angle and it invents the rest of the head from statistical averages rather than from your character.
- Style bleed. Background style, grading, film grain and effects all influence how the subject is rendered, especially in stylised output where texture is doing a lot of the identity work.
- Motion ambiguity. Vague verbs like "walks" or "reacts" leave the model free to invent gait, posture and gesture habits that were never part of your design.
The crucial insight is that none of these drivers is fixed by writing a longer prompt. Longer prompts usually make dilution worse. They are fixed by architecture: deciding in advance which traits are immutable, which are negotiable, and which are forbidden.
The Lego Pixel Idea in Plain Terms
The name welds two ideas together. Lego means modularity — a character is built from parts that can be swapped, re-described and recombined without rebuilding the whole. Pixel means quantisation — deliberately reducing detail into a low-resolution mosaic so that small generative deviations no longer change who the character is.
Think about pixel art. A 16x16 sprite survives compression, scaling, palette swaps and re-rendering because its identity lives in a handful of high-contrast features, not in fine texture. There is simply nothing subtle enough to drift. Apply that logic to video and your character anchors to roughly eight invariant traits: silhouette shape, hair mass, skin tone band, a signature garment colour, one accessory, default posture, and two or three facial landmarks.
Everything else is free to vary. The colour of a shirt that only appears in one scene can wobble. The exact weave of a fabric can change between shots. What cannot wobble is the set of anchors that lets a viewer recognise the same person at thumbnail size.
The payoff is robustness. When the model has freedom in the details but strict constraints on the anchors, drift becomes cosmetic rather than catastrophic. A wrong cuff is a nuisance. A wrong face is a re-render of the entire series.
Building a Character Block Manifest
A manifest is a short structured document — effectively a character sheet in plain text — that every prompt is built from. Keep it in the same file as your shot list so prompts and continuity never diverge. If a shot prompt does not reference the manifest, it is not part of the series; it is a test render.
A workable block set looks like this:
- Silhouette block — height, build, shoulder line, head-to-body ratio.
- Head block — hair mass, hairline, hair length, facial hair, headwear.
- Face block — eye shape and colour, brow weight, nose profile, jaw, plus two or three distinguishing marks such as a scar, a mole or a chipped tooth.
- Palette block — three to five colour values that never change, ideally written as hex codes so nobody argues about what "dusty teal" means.
- Wardrobe block — the signature garment, its cut, its fabric and its fastenings.
- Prop block — anything the character always carries, plus which hand holds it.
- Motion block — gait, gesture habits, resting posture and default speed.
Invariants, Variables and Forbidden Moves
For each block, write three lists. This three-list structure is the real engine of the method, because it converts taste into rules that survive handoffs between sessions and between people.
Invariants never change across the series. For a courier character, that might be: 172 cm, narrow shoulders, close-cropped dark hair with a high hairline, olive skin band, bottle-green jacket, silver thumb ring on the right hand, walks with a slight forward lean.
Variables may change per scene without breaking identity: jacket open or closed, sleeves rolled, hair damp from rain, ring turned inward, a temporary bandage on one forearm.
Forbidden moves break identity instantly, and they are the list most teams forget to write. Changing eye colour. Removing the signature prop. Switching from a green palette family into a red one. Lengthening the hair beyond the collar. Making the character taller for a hero shot because the composition looks better. Every one of these has ruined a shot sequence at some point.
Casting the Character in Roughly 140 Words
Write a casting paragraph containing only identity information — no scene, no camera, no lighting, no mood. Concrete nouns beat adjectives every time: "short boxy denim jacket with copper buttons and a frayed hem" outperforms "cool retro jacket". End the paragraph with two or three "never" statements lifted directly from your forbidden list.
Keep that paragraph frozen. Paraphrasing it later, even helpfully, invites drift.
Creating a Canonical Turnaround and Slicing It Into Reference Crops
Quality at this stage pays off for the whole project. A muddy turnaround will haunt every later shot, because every later shot is conditioned on it.
Step 1: Generate the Turnaround
Produce a clean image set in neutral light on a plain background: full-body front, three-quarter, profile, and a back view. Add a posing sheet with four or five common stances. Fix the seed and the palette before you iterate, then change only the blocks you dislike rather than re-rolling the whole character and hoping.
Step 2: Slice the Design Into Blocks
Take the turnaround and literally extract the parts. Crop the head, the torso, the hands with the prop, the boots. Save each as its own reference file with a clear name such as character_a_head_3q.png. In practice this gives you a library of five to eight files per character instead of a single hero image — and a library is what reference conditioning actually wants.
Step 3: Test the Library Before You Trust It
Run a five-shot consistency test: one close-up, one medium, one wide, one from behind, one in motion. Do not move to full production until that test holds. Budget the afternoon; it is far cheaper than re-rendering a third of the series.
Writing Shot Prompts That Resist Drift
The Five-Part Shot Prompt Formula
A stable shot prompt has five parts, always in the same order:
- Identity anchor line, pulled verbatim from the manifest.
- Action and emotion — one clear verb, one clear emotional register.
- Camera and framing — distance, height, lens feel, movement.
- Lighting and environment — source, direction, quality, time of day.
- Palette reminder plus negative constraints — the three to five values, then what must not appear.
Keeping the identity line word-for-word identical in every shot is the single highest-leverage habit in the entire workflow. It feels repetitive. That repetition is the point.
Ordering and Weighting References
Most modern tools accept several reference images or a dedicated character slot. Feed them in a deliberate order: full-body silhouette first, then face close-up, then prop detail. Some tools weight the first reference most strongly, others blend evenly, and a few re-weight based on aspect ratio.
Run a five-shot test with shuffled order and keep whichever sequence produces the most stable face. Write the winning order into the manifest so nobody has to rediscover it.
Negative Constraints That Actually Help
Negatives work best when they are specific and physical: "no beard", "no long hair", "no red garments", "no wide-angle distortion on the face". Broad negatives like "no bad anatomy" do almost nothing because they describe a quality judgement rather than a visual fact.
Adapting the Method Across Generator Families
Different model families reward different parts of the method. This is not about crowning a winner; it is about matching technique to tool.
| Model family | Strength | How to apply the method |
|---|---|---|
| Image generators for stills and sheets | High-fidelity turnarounds and crops | Build the turnaround and block crops here first |
| Text-to-video suites with motion and camera control | Smooth movement, strong camera language | Use image-to-video from a reference frame, keep identity lines short |
| Stylised animation tools with preset looks | Expressive motion, graphic aesthetics | Lean on the palette block; expect the face block to need re-anchoring |
| Long-take frame pipelines | Fewer cuts, continuous action | Re-inject block references at intervals, not only at frame one |
A few practical notes that apply across all of them. Generate at the model's native resolution instead of upscaling later, because upscaling amplifies small identity errors. Keep the aspect ratio fixed across a series; changing from vertical to horizontal mid-project changes framing and therefore changes how faces are rendered. And always test a new model against your existing manifest rather than rewriting the manifest to suit the model — the character should outlive the tool.
For stylised output, pixel-block logic works best because the model's own aesthetic already discounts fine detail. For photoreal intimacy — tight close-ups with micro-expression — blocks must be tighter and tolerance smaller, which raises the cost of every iteration and makes the five-shot test more important, not less.
An End-to-End Production Workflow
Pre-Production
Lock the script and shot list first, then build manifests, then the turnaround, then the consistency test. Working in the other order — generating pretty shots before you know what the series needs — produces beautiful footage you cannot reuse.
Batch Similar Shots
Generate in batches of similar shots: same location, same lighting, same camera distance. Batching reduces the number of variables changing between generations, which makes it obvious which element is causing drift when it appears. If three shots in a row look right and the fourth looks wrong, the fourth introduced a new variable, and you can find it.
Log Every Generation
Save every prompt beside its output, with seed, reference set and model version. This is tedious for the first ten shots and invaluable for the ninetieth. It is also the only way to reproduce a happy accident you did not plan for.
Edit Before You Polish
Cut the first assembly before fixing individual shots. Some drift disappears in the edit; a clip that looks wrong in isolation may read perfectly at two seconds between two other shots. Fix only the offenders that survive the cut, and fix them by block: palette problems get a palette fix, not a whole new prompt.
Continuity Documents, Versioning and Handoffs
Track per character: manifest version, reference file names, approved seeds and a short change log. When a collaborator opens the project months later, the document should explain every look decision without requiring a conversation.
Version the manifest properly. When you change an invariant — say, you decide the character's hair is now shoulder length — publish it as a new version and note which shots were generated under the old one. Silent edits destroy continuity faster than any model quirk, because half the series will follow a rule nobody wrote down.
A useful habit is a one-page style bible that sits above the manifests: overall palette, grading approach, camera grammar, and a short paragraph on tone. It keeps two characters from the same series feeling like they came from different projects.
Failure Modes, Fixes and When Not to Use the Method
The Five Failure Modes Worth Knowing
- Face drift. Usually prompt dilution. Shorten the shot prompt and move the identity anchors to the very front.
- Wardrobe creep. Caused by vague garment language. Name the fabric, cut and fastening; ban synonyms for the signature item.
- Palette slide. Often introduced by lighting or grading instructions. Restate the palette after the lighting line, never before.
- Style bleed. Background style contaminating the character. Separate style from subject with an explicit boundary sentence.
- Identity lock. Over-constrained prompts produce mannequins. Loosen the motion block and let expression vary; a character who cannot smile is not consistent, just frozen.
Most projects fail at the fourth or fifth shot, not the first. The first shot is easy because there is nothing to contradict. That is exactly why a five-shot test before full production is worth the time.
When This Method Is the Wrong Tool
The Lego pixel approach is a consistency tool, not a universal aesthetic. It is the wrong choice for one-off flagship shots where you can cherry-pick the best of fifty attempts, for documentary-style footage where imperfection reads as authenticity, and for characters whose appeal depends on fine photoreal detail at very close range, where quantising features costs more than it saves.
It also adds friction to genuinely short projects. If you need three shots total, write three good prompts and move on. The method earns its keep when you need twenty or more shots of the same person, across sessions, with more than one person touching the project.
FAQ
Do I need a specific tool to use this method?
No. It is a workflow pattern, not a feature. Any image generator plus any image-to-video or text-to-video model can support it. The method lives in how you describe, reference and verify characters — not in a button.
How many reference images should each character have?
Three to six is the practical sweet spot: silhouette, face, prop, plus one or two angle references. More references can confuse models that blend weights evenly, producing an averaged face that resembles nobody.
Can I use it for two characters in the same shot?
Yes, but keep each character's blocks isolated in the prompt and describe their spatial relationship explicitly. Two characters in frame roughly doubles drift risk, so budget extra test renders and expect to fix one of the two more often.
What about voice and dialogue?
Treat voice as a block. Fix pitch range, pace and vocabulary habits, then describe them in the manifest so performance stays aligned with visual identity. A character whose voice changes between scenes reads as a different person even if the face holds.
How long should each shot be?
Two to four seconds gives the most control. Longer takes accumulate drift because the model has more frames in which to reinterpret the character. If a beat genuinely needs eight seconds, split it into two shots at a natural cut and keep the camera close across the join.
How do I know when a project is consistent enough?
Watch the first cut as a viewer, not as a prompt engineer. If you stop noticing the character and start following the story, consistency has done its job. Chasing pixel-perfect uniformity past that point costs time without improving the film.
Does the method work for stylised, non-human characters?
It works especially well, because stylised designs already compress identity into a few strong shapes. A robot or a creature defined by silhouette, palette and one signature detail is exactly the kind of character that survives aggressive quantisation.
What is the fastest way to fix a series that already drifted?
Do not re-render everything. Pick the shot with the strongest face, treat it as the new canonical reference, rebuild the block crops from it, and regenerate only the shots where the identity breaks. Then write the manifest retroactively so the same mistake cannot repeat.


