Offre à Durée Limitée : 50% DE RÉDUCTION sur votre premier mois de Pro & Ultra 🎉

Pixel Lego Method for Consistent AI Video Characters

Sep 13, 2026

Consistent characters are the hardest problem in AI video. Camera moves, lighting shifts, and scene changes are all solvable with a good prompt, but the moment a story needs the same face in more than one shot, most workflows fall apart. The eyes drift, the jawline changes, the jacket becomes a different jacket. Pixel Lego is the name many Thai-speaking creators use for the technique that fixes this: building character identity out of small, locked image tiles and then fusing those tiles back together for every shot.

Rather than hunting one perfect reference image, you author a small library of tiles — a front view, a three-quarter view, a profile, an expression grid, a wardrobe plate — and then fuse them into a single identity reference before every generation. The tiles behave like the blocks of a construction toy. Each one is small, flat, and interchangeable, but the assembled model is stable enough to survive a hundred shots.

This guide covers why the approach works, how the underlying processing is modular, how to control keyframes and emotional continuity, and where the workflow breaks. It is written as a practical playbook. You can follow it with any image and video model that supports multi-reference input.

Why character drift breaks AI video projects

Text-to-video models do not have memory. Every generation starts fresh from noise, guided by a prompt and, if you are lucky, one or two reference images. When you run a second shot of the same scene, the model re-invents your character from scratch. Small differences compound, and by shot five the protagonist looks like a cousin rather than the same person.

Character drift shows up in predictable places:

  • Facial geometry: eye spacing, nose width, and jaw angle shift slightly between clips.
  • Age cues: skin texture and facial fullness drift younger or older depending on lighting.
  • Wardrobe: logos, stitch lines, and colors mutate, especially in motion.
  • Hair: parting, length, and highlight patterns change with camera angle.
  • Proportion: head-to-body ratio wanders when the framing changes from wide to close.

Any one of these is forgivable in a still. Together, across ten clips, they destroy the illusion that a single person exists in your story. Viewers cannot articulate the problem, but they feel it as cheapness.

The practical consequence is rework. Creators regenerate shots hoping the model will land closer, burn hours on cherry-picking, and still end up with a sequence that cuts badly. Pixel Lego reduces this by removing chance. Instead of asking the model to guess who the character is, you hand it a fixed, modular identity built once and reused everywhere.

What Pixel Lego actually is

Think of it as a construction system for identity rather than a single prompt trick. The method has four moving parts.

First, tiles. Each tile is a tightly cropped, well-lit image showing one attribute: the face front-on, the same face turned, a hand, a jacket, a pair of shoes, a signature accessory.

Second, an assembly step. Tiles are fused into one composite reference sheet, usually through multi-image fusion or reference-guided image generation, so the model sees a single coherent subject instead of a pile of disconnected photos.

Third, a lock. The assembled sheet, plus a written identity description, becomes the canonical character definition that never changes mid-project.

Fourth, reuse. Every shot prompt cites the same locked definition. Only action, camera, and lighting change.

The name is a metaphor for the important property: modularity. Because each tile isolates one attribute, you can repair one block without rebuilding the character. If the jacket is wrong, you swap the wardrobe tile and reassemble. If the character ages across a story, you swap one face tile and keep everything else.

Modular image processing: how the fusion pipeline works

The pipeline below is model-agnostic. It works whether you generate images with a diffusion model, an image-editing model, or a video model that accepts reference images.

Step 1 — Build a tile sheet

Shoot or generate six to ten tiles of the same character on a neutral background. A dependable starter set:

  • Front face, neutral expression, even lighting.
  • Three-quarter left and three-quarter right.
  • Profile.
  • Expression variations: neutral, smile, concern, anger.
  • Full-body front, for proportion and wardrobe.
  • Detail tile: hands, or a defining accessory.

Generate the tiles with the same seed family and the same lighting description so that color temperature stays constant. Inconsistent white balance between tiles is the most common cause of a fused identity that looks patchy.

Step 2 — Fuse the tiles into one identity reference

Multi-image fusion means feeding several tiles into a single generation and asking for one subject rendered from the angle you need. Prompt it explicitly as an identity merge, not a collage. A workable phrasing pattern:

"A single character identity reference: the same person as the attached views, front-facing, neutral studio lighting, plain grey background, sharp detail on eyes, nose and jawline."

The output should look like one photo, not a grid. If you get a grid, the model treated the input as a layout, so reduce the number of tiles and repeat the phrase "one person, one photo".

Step 3 — Write the identity card

The identity card is a short paragraph of prose that describes what the tiles cannot: face shape, eye color, hair texture, age range, build, and recurring wardrobe. Keep it under 90 words and identical in every prompt. Prompts that drift in wording produce images that drift in appearance.

Step 4 — Generate, inspect, repair the weak tile

After each generation, check four things: eye spacing, hairline, jaw angle, and wardrobe color. Whichever mismatches, fix that tile only. Reassemble. Never patch a whole character because one attribute is off.

Step 5 — Lock and version

Save the tile set and identity card as a versioned folder, for example character-name_v3. When you deliberately evolve the character, create a new version rather than editing the old one, so earlier shots remain reproducible.

Keyframe control and continuity between shots

Even with a perfect identity, a sequence fails if motion and framing do not connect. Keyframe control is the discipline of defining the start, middle, and end of each clip so the cut is invisible.

A reliable method is to define three frames per clip: an entry frame that matches the previous clip's exit, a midpoint that establishes the new beat, and an exit frame that sets up the next shot. Generate the entry and exit frames as images using the locked identity, then animate between them with image-to-video. This gives you continuity that pure text-to-video cannot produce, because the model no longer invents the beginning and end of the motion.

Three continuity rules matter more than any others:

  • Match the light direction. Track the key light position as a line in your shot list and keep it consistent within a scene, even when the camera moves.
  • Match the lens. If a scene is shot at a 35mm equivalent, keep it there. Jumping from wide to telephoto between consecutive shots reads as a different production.
  • Match the exit pose. The final frame of clip A and the first frame of clip B should share body orientation and eyeline.

A companion technique is object anchoring: keep one visible prop consistent, such as a mug, a bag, or a vehicle, so the viewer's eye has a fixed reference across cuts. Anchors make identity drift far less noticeable and, more importantly, less damaging.

Preventing emotional drift across the sequence

Emotion is where most identity systems quietly fail. A face that is recognizably the same person can still read as a stranger if the emotional register jumps between clips. A character who is quietly worried in shot one should not become theatrically alarmed in shot two unless the story earns it.

Build an emotional continuity ladder, one line per shot, using a simple intensity scale such as 1 to 5. Then translate the level into prompt language rather than adjectives alone:

  • Level 1: stillness, lowered eyelids, mouth relaxed, shoulders level.
  • Level 2: slight brow tension, lips together, weight shifted to one foot.
  • Level 3: clear brow furrow, jaw set, hands occupied.
  • Level 4: eyes wide, breathing visible at the shoulders, forward lean.
  • Level 5: full body engagement, motion blur in the frame.

Add this ladder to your identity card as a separate field so it stays stable across the project. When you need a jump in intensity, insert a bridging shot — a cutaway, a reaction insert, or a beat of stillness — rather than asking the model to leap two levels in one clip.

Emotional continuity also depends on pacing. If you cut faster as intensity rises and slower as it falls, viewers accept a wider range of emotional variation. If you keep a constant rhythm, small mismatches become obvious.

A shot-by-shot workflow you can copy

This is the working loop that keeps a project consistent from the first frame to the last.

  1. Write the shot list with columns for shot number, action, camera, lighting key, lens, emotion level, and required tiles.
  2. Assemble and lock the identity reference for the protagonist, and repeat for any recurring secondary character.
  3. Generate entry and exit frames for shot one using the locked identity.
  4. Animate shot one with image-to-video, using the identity reference as visual guidance.
  5. Inspect against the continuity checklist: eyes, hairline, jaw, wardrobe, light direction, lens.
  6. Repair the weakest tile if any attribute drifted, then regenerate only the failing clip.
  7. Copy the exit frame and use it as the entry frame for shot two.
  8. Repeat until the sequence is complete, then assemble and review at full speed, not frame by frame.

Reviewing at full speed matters because consistency is a perceptual property, not a pixel-level one. A tiny mismatch invisible at 24 frames per second is irrelevant; a large one that only appears when paused is usually safe to ship.

Choosing tools for this workflow

You do not need one tool that does everything. You need a stack where each layer is replaceable.

  • Tile generation: an image model with strong identity preservation and reliable seed control.
  • Fusion: an image-editing or reference-guided model that accepts multiple images in one request and can output a single coherent subject.
  • Animation: an image-to-video model that accepts a start frame and, ideally, an end frame.
  • Continuity review: a simple side-by-side contact sheet tool. A 4x4 grid of first and last frames per shot catches more drift than any automated score.
  • Asset management: plain folders with versioned names. Fancy asset systems are optional; disciplined naming is not.

Decision criteria when comparing options: does it accept multiple references in one call, does it honor an explicit end frame, does it preserve material detail in motion, and can you reproduce a render from the same inputs a week later. Reproducibility outranks raw quality, because a slightly weaker model you can re-run beats a stronger one you cannot.

Common failure modes and how to fix them

These are the problems you will actually hit, in the order most creators hit them.

Patchwork identity. The fused reference looks like several people blended. Cause: inconsistent lighting or background between tiles. Fix: regenerate tiles with identical lighting language, then fuse again with fewer inputs.

Grid output. The model renders a collage. Cause: too many references at once, or prompt language implying a layout. Fix: cut to three or four tiles, and state "one photo of one person" explicitly.

Wardrobe drift in motion. Colors and logos mutate mid-clip. Cause: fabric detail is under-specified and the model improvises. Fix: add a wardrobe tile and describe material, not just color.

Sudden age shift. The character looks older after a lighting change. Cause: harsh shadows read as wrinkles. Fix: keep lighting descriptions consistent across the scene, or deliberately change them scene-wide rather than per shot.

Expression whiplash. Emotional jumps feel unmotivated. Cause: intensity not tracked per shot. Fix: use the ladder and insert bridging shots.

Identity collapse on wide shots. Character loses specificity as the camera pulls back. Cause: too few pixels on the face. Fix: accept that wide shots carry identity through silhouette, hair, and wardrobe, and lock those three explicitly.

Where this approach does not help, and a starting drill

Pixel Lego solves a specific problem: keeping one character recognizable across many shots. It is not a cure for weak storytelling, and it does not fix a model that cannot hold anatomy in fast motion. Long dialogue scenes with heavy facial performance still benefit from keeping shots short and cutting on reaction beats.

For non-human characters, the same logic applies with one change: tiles capture design language rather than facial identity, so silhouette and material detail carry more weight than eye spacing.

And for brand work, treat the identity card as a legal document. Lock it, version it, and get approval on the tile set before generating a hundred clips you may have to throw away.

If you want to test the method today, run this drill in under an hour. Pick one character. Generate four tiles. Fuse them. Write a 60-word identity card. Produce three shots: a walk, a turn, and a close-up. Then put them side by side.

If the character survives all three shots with the same eyes, hairline, and jacket, your tile set is working and you can scale the project. If not, the tile that failed is visible in the contact sheet — repair it and run the drill again. Consistency is not something you hope for in AI video. It is something you build, one block at a time.

FAQ

How many tiles does a character need?

Four is enough to start, six to ten is comfortable for a full production, and more than twelve usually creates fusion problems rather than solving them.

Should I generate tiles with the same seed?

Yes, where the model exposes it. Keeping the seed family close reduces lighting and color variation between tiles, which is the primary cause of patchwork identity.

Can I reuse one identity reference across different scenes?

Yes, and you should. The identity reference is scene-independent. Only the lighting, lens, and emotion fields in your prompt should change from scene to scene.

How do I handle a character who ages during the story?

Version the character. Create an older tile set, keep the same identity card structure, and regenerate the exit frames. Never blend ages inside one version, because the model will average them.

What do I do when only one clip drifts?

Regenerate that clip alone using the neighboring clip's exit frame as the entry frame. If it drifts a second time, the identity reference is the problem, not the prompt.

Is this workflow only useful for realistic characters?

No. Stylized, animated, and illustrated characters benefit from the same modular discipline, with silhouette and line quality replacing facial geometry in the tile set.

How many shots can one identity reference sustain?

With a locked reference and entry-frame chaining, a strong project can run thirty to fifty shots before any noticeable creep, and contact-sheet reviews catch the slow drift before it becomes visible to an audience.

Alexander

Alexander