Why AI Video Still Breaks Between Shots
Ask any generation engine for a rain-soaked alley at dusk and you get something that reads like a film still. Ask for a second angle of the same alley and the trouble starts. The character's jaw softens, the jacket drifts from olive to grey, a window appears where there was brickwork. By the sixth shot you are not cutting a scene, you are assembling a collage of near-misses that each look good alone and wrong together.
The reason is structural rather than aesthetic. Most engines treat each generation as an independent event. When you hand one a reference image, it compresses that image into a loose style signal instead of a geometric constraint. The model borrows mood, not measurements. Facial structure survives only approximately, and the error compounds with every new frame. Block pixel referencing, sometimes described informally as a Lego pixel approach, exists to close exactly that gap.
Instead of asking a model to remember a character, you give it a grid of small, tightly cropped reference tiles. Each tile is locked to one region or property: eyes, hands, fabric weave, storefront lettering, the falloff of light on a wall. Every tile is a fixed unit that can be recombined like a building block. The model no longer has to guess what consistency means, because it is being told where to look and what to preserve.
The rest of this guide covers what the technique actually changes, how to assemble a grid that survives a full sequence, how to prompt against it, how different classes of engine respond to it, and how to fold the whole method into a pipeline a small team can repeat without rebuilding it for every project.
What Block Pixel Referencing Actually Means
Strip away the jargon and this is a machine-readable version of something animation and visual effects artists have kept for decades: a shot bible. In hand-drawn animation, a character sheet pins down the exact proportions of a face from every angle. Block referencing does the same job in a format a diffusion model or a video engine can consume as conditioning input.
The technique has three moving parts. First, a set of tightly scoped crops. Second, a layout that gives each crop a spatial address in the conditioning signal. Third, a locked configuration that prevents the operator from accidentally introducing variation between runs. If any of those three is missing, quality regresses toward the single-reference baseline.
The reference pixel unit
A reference pixel unit is a small, high-detail crop that satisfies three conditions at once. It is framed around exactly one visual property, not a whole scene. It is reproduced at a resolution high enough that the model can extract texture rather than just silhouette. And it is labeled, so that both you and the model know what it governs.
A workable set for one recurring character usually includes:
- A neutral front-face crop framed at eye and nose level
- A three-quarter crop that captures cheekbone, jaw, and ear structure
- A hand crop with fingers separated, well lit, and free of overlap
- A fabric crop showing weave, stitching, and how shadow behaves in the folds
- A silhouette crop against a strongly contrasting background
- A color card carrying the three dominant wardrobe values, sampled not guessed
The logic here is redundancy through specificity. One portrait reference forces the model to generalize everything from a single sample. Six tightly scoped crops force it to reproduce measurable detail at six specific locations.
Why blocks beat one hero frame
A single reference image gets globally averaged. When the model diffuses a new frame, it blends your reference into one broad style vector. Facial structure survives loosely, and drift compounds across a sequence. Block references give the conditioning signal a spatial address: the eye crop influences the eye region, the hand crop influences the hand region, and errors stay local instead of spreading across the whole frame.
This is also why block referencing pairs so naturally with image-to-video work. You are not asking the engine to invent continuity. You are asking it to preserve geometry you have already approved.
Where the grid sits in a production pipeline
The order matters. Approvals come first, then references, then prompts, then generation, then assembly, then finishing. A grid built before the character design is signed off will be rebuilt twice. A grid built after the first three shots are generated is a repair job, not a foundation.
The Three Consistency Failures Blocks Fix
Character, environment, and grade fail in different ways. Treating them as one problem is why so many sequences still fall apart even when the references look clean.
Identity drift
Identity drift is the slow mutation of a character across shots. It shows up first in the eyes, then the hairline, then the jaw. By the fifth generation your lead looks like a cousin of the person you approved. The cause is usually procedural: each new frame is generated relative to the previous output rather than the original reference, and errors accumulate the way photocopies of photocopies degrade. Block references break the chain by conditioning every frame on the original crops.
Environment drift
Locations are harder than faces because they contain more independent objects. A cafe has windows, chairs, signage, condensation on glass, and a specific quality of afternoon light. If the engine regenerates the space from scratch each shot, the geometry shifts. Fixed blocks for the three or four objects a viewer will actually notice, such as the sign, the window frame, and the counter edge, anchor the set without over-constraining everything else.
Grade and texture drift
Even when characters and sets hold, the grade wanders. Warm scenes cool down, contrast creeps upward, grain texture changes character between cuts. A small palette block plus a tonal reference crop stabilizes the look so that cuts feel like one continuous piece of photography rather than a stock-footage assembly.
Building a Reference Grid Step by Step
The build itself takes an afternoon. Keeping it clean over a month of iteration is the real work, and that is where naming and locking pay off.
Step one: normalize the source material
Pull every approved frame into one folder. Crop to your target aspect ratio and resample to a consistent short edge, using 1024 pixels as a comfortable floor for most current engines. Apply identical color management so no reference is accidentally warmer than another. If you are starting from photography, clean compression artifacts before cropping, because block references amplify whatever noise they inherit.
Step two: tile and label the sheet
Lay the crops out in a fixed grid, usually three columns by three rows or four by two. Keep consistent gutters between tiles, because some engines read adjacent tiles as one image when they touch. A clean 3x3 sheet gives you nine anchors: face, three-quarter, eyes, hands, wardrobe, silhouette, environment hero, palette, and lighting reference.
Labeling is not bureaucracy. When you iterate across weeks, an unlabeled sheet becomes unusable, and a collaborator cannot tell a v2 crop from a v7 crop. Name tiles by function and revision, for example lead_face-neutral_v3 or cafe-window-frame_v2. Keep the label short enough to survive a filename limit and specific enough that nobody has to open the file to understand it.
Step three: test two input patterns per engine
Not every engine accepts the same pattern. Some want one composited sheet, others accept multiple reference images with relative weights. Test both variants and compare the first output against your reference at 200 percent zoom. Inspect the eye region first, because it is the most punishing area and the fastest way to see whether spatial addressing is actually happening.
Step four: lock the configuration
Once a setup reproduces the reference faithfully, freeze it. Record the engine, checkpoint, guidance value, seed behavior, reference order, and any motion settings. Iteration is where consistency dies quietly: a small convenience change in step order can break continuity three shots later, and you will blame the model instead of the process.
Step five: version the assets, not just the project
Keep a reference sheet per character, per location, and per grade. When a wardrobe change happens in episode four, branch the character sheet rather than editing it in place, so earlier shots remain reproducible. This one habit saves more time than any prompt trick.
Prompting Against a Reference Grid
When references carry the visual load, the prompt carries the structural and narrative load. Getting that division right is the single biggest quality lever available to you.
Describe structure, not mood
Replace adjectives about appearance with instructions about action, framing, and motion. Instead of writing a cinematic portrait of a woman with striking green eyes, write medium close-up, subject turns from left to right, slow dolly in, practical light from the window behind the camera. The reference already defines the eyes. The prompt should define what happens next.
Keep negative constraints short
Long negative lists flatten output. Pick the two or three failure modes you are actually seeing in this sequence, such as face warping, extra fingers, or background duplication, and name only those. Re-evaluate the list every few generations instead of carrying dead constraints forward out of habit.
Order references by priority
Where an engine accepts ranked references, the first slot carries disproportionate weight. Put the crop governing the most drift-prone feature in slot one. For character work that is almost always the face. For product work it is usually the logo or label, because fine typography degrades faster than anything else in the frame.
Write prompts as reusable templates
A template with variable slots beats a fresh paragraph every time. Something like: shot size, subject action, camera move, lighting source, mood word, duration. Fill the slots, keep the grammar stable, and you remove an entire category of accidental variation from your sequence.
Choosing the Right Setup for Different Engine Classes
Engines respond to block referencing in recognizably different ways. Knowing the pattern saves hours of trial and error, and it prevents the classic mistake of applying a text-to-video reference strategy to an upscaler.
| Engine class | What works best | Watch out for |
|---|---|---|
| Text-to-video with multi-reference input | Flattened 3x3 sheet, ranked references | Weak environment fidelity relative to character fidelity |
| Image-to-video | First frame plus two anchor blocks | Drift during fast motion |
| Keyframe-first tools | Exactly two blocks: one identity, one environment | Diluted conditioning when flooded with nine tiles |
| Image upscalers and refiners | Palette and grain blocks only | Plastic skin when fed facial crops |
| Stylized animation engines | Silhouette plus palette blocks | Over-constrained motion and stiff results |
A useful rule of thumb: match the granularity of your references to the granularity of the engine's output control. Coarse engines want coarse references. An engine that barely respects your prompt will not respect nine tiles either, and adding more inputs only muddies the conditioning.
Decision criteria in practice:
- If identity holds but the set wobbles, add environment blocks, not more face blocks.
- If the set holds but skin looks synthetic, reduce local detail references and let the grade block lead.
- If motion smears, shorten the action described in the prompt before touching the references.
- If the first frame is wrong, fix it with image-to-image before you generate any video at all.
Environment, Camera, and Grade Discipline
Locations fail differently from characters. Rather than mutating organically, they lose objects. A chair quietly disappears between shots, a doorway moves two meters left, a window changes proportions. Viewers may not identify the error consciously, but the scene stops feeling like a place they have been.
Build a location sheet with the same discipline you use for characters. Choose four anchor objects that a viewer will notice plus one lighting reference. For each new shot, include the two most relevant anchors and the lighting block. In a kitchen scene that might be the window frame and the countertop edge. On a street it might be a storefront sign and a curb line.
Add a rough overhead diagram to your production notes. Even a sketch prevents the most common spatial continuity error in generated scenes: characters who enter from the wrong side or stand in physically impossible positions relative to established geometry. It also speeds up blocking decisions when you are generating ten shots in an afternoon.
Camera language needs the same restraint. If you want a sequence to read as shot on one lens, generate your lighting and palette blocks from frames that already exhibit that lens character, with its depth of field, bokeh shape, and flare behavior. Then keep three parameters steady: focal length, depth-of-field character, and the amount of visible grain. Mixing wide and long-lens language across a dialogue scene produces an unnatural jump even when the character holds perfectly.
A practical test: generate five shots with your grid, cut them together with sound muted, and watch. If the sequence reads as one shoot, your camera language is consistent. If it feels like footage from five different sources, your references are too heterogeneous to be doing their job.
Finishing: From Generation to Delivery
Structural consistency and perceived sharpness are different problems solved at different stages. Structural consistency comes from the reference grid and prompt discipline. Perceived sharpness comes from delivery-side processing: careful scaling, controlled sharpening, and grain that matches the intended output. Trying to fix drift with an upscaler produces smooth, waxy frames that lose the texture that made the original appealing.
A finishing order that holds up in practice:
- Generate at native resolution with references locked and configuration unchanged.
- Select the best take per shot, judging motion and expression rather than the sharpest single frame.
- Repair small artifacts locally before any global processing, because global passes make local fixes harder.
- Upscale in one pass with a mild detail model rather than three aggressive ones.
- Add grain and halation last, at final resolution, so the texture is consistent across cuts.
- Grade once, after the edit is locked, so color decisions are made against the real cut.
If your deliverable is vertical social video, reframe after the edit rather than before. Cropping references to a vertical ratio too early changes how the engine interprets spacing and composition, and it quietly degrades your master assets for any future widescreen version.
Keep a short delivery note with each sequence: resolution, aspect ratio, frame rate, grade notes, and which reference sheet revision was used. Six months later, that note is the only thing that lets you regenerate a single shot without rebuilding the entire look.
Common Mistakes and How to Fix Them
Using too many references. Nine tiles is a ceiling, not a target. If the engine starts ignoring inputs, cut back to three and rebuild one block at a time until you find the tile that helped.
Mixing sources from different sessions. Different lens, light, and grade inside one grid teaches inconsistency. Regenerate references from a single controlled session whenever possible, even if that means reshooting a still frame.
Leaving text in the references. Incidental signage, watermarks, or interface elements can be reproduced as if they were part of the subject. Crop clean, and check the edges of every tile before you lock the sheet.
Chaining generations. Feeding output back in as a reference is the fastest route to drift. Always return to the original sheet, even when the previous frame looked better.
Ignoring motion blur. References are static; motion is not. Prompts that describe fast action with hard cuts often produce smeared faces. Slow the action or reduce camera movement and let the reference do its work.
Over-sharpening after generation. This creates halos that look worse in motion than the softness you were trying to correct.
Skipping version control. Without a naming convention, a project with forty reference sheets becomes unmaintainable within a week, and collaborators start guessing which file is current.
Fixing the wrong problem. When a shot looks wrong, decide first whether the failure is identity, environment, grade, or motion. Each has a different fix, and reaching for an upscaler is the wrong answer for all four.
FAQ
Does block referencing replace prompt writing?
No, it redistributes the work. References handle appearance, prompts handle action and framing. When those two responsibilities overlap, results get muddy and tweaking becomes guesswork.
How many reference tiles are too many?
Past seven or eight, most engines begin averaging rather than honoring individual tiles. Start with three, then add a tile only when you can name the specific failure it fixes.
Can I reuse one sheet for both images and video?
Yes, but the video pass usually needs a tighter crop set. Video engines spend more of their conditioning budget on motion, so references should be more redundant and more consistent between tiles.
Why does consistency hold for two shots and break on the third?
Usually reference ordering or seed handling changed between runs, or a frame was chained from previous output. Verify that your inputs match the locked configuration before you blame the engine.
Should references be photographs or generated frames?
Photographs carry more texture and less model bias, while generated frames match your target style more closely. A hybrid works well: photographic references for structure, one generated frame for grade and texture.
How do I handle props that appear only once?
Do not add them to the master sheet. Generate a single-purpose reference for that shot, use it, and discard it so the master sheet stays stable across the sequence.
What about audio consistency?
It is a separate discipline, but the logic is identical. Define a small set of anchors, such as room tone, a voice reference, and an ambience bed, then reuse them instead of regenerating per scene.
Do I need a powerful machine to work this way?
The grid itself is lightweight. What costs time is iteration, so the practical benefit of locking a configuration is that you stop burning hours on re-tests you already proved unnecessary.
Getting Started This Week
Pick one short sequence: three to five shots, one location, one recurring character. Build a 3x3 reference sheet from the best frames you have, label every tile, and lock an engine configuration before you generate anything. Then resist the urge to change settings mid-run, even when a single frame tempts you to.
The measurable result you are looking for is simple. When you cut the finished shots together, a viewer should not be able to point at the moment continuity broke. If they cannot, the block pixel method has done its job, and everything you built along the way becomes reusable infrastructure for the next project rather than a one-off experiment.




