Limited Time Offer: Get 50% OFF your first month of Pro & Ultra plans 🎉

AI Video Character Consistency: Heights, Sets, and Scale

Sep 16, 2026

Why Character Scale and Set Realism Break AI Video

Generated video looks convincing right up until a person walks through a doorway that is clearly too short, sits on a stool that seems to swallow them, or appears three inches taller in the next shot. Those errors are not aesthetic nitpicks. They destroy the viewer's trust instantly, and they are the most common reason a promising AI-generated scene gets rejected during review.

The root cause is almost always the same: creators describe characters with adjectives and describe sets with moods. "Tall man in a leather jacket" and "cramped backstage dressing room" leave the model to invent scale from nothing in every single shot. Because diffusion-based video models sample from probability distributions rather than a persistent 3D scene graph, the invented scale drifts from frame to frame and from cut to cut.

Fixing this means moving from description to specification. You need a compact set of numbers, a reference board, and a continuity document that travels with you from the first prompt to the final edit. This guide covers that system end to end: how to measure human proportions in a way a model can actually use, how to lock a look across shots, how to design tight interior spaces that read correctly, and how to structure a repeatable production workflow around the whole thing.

The Measurement Layer: Converting Human Proportions Into Usable References

Before you touch a prompt, decide what your character physically is. This is the foundation everything else rests on, and it takes about thirty minutes to set up properly.

Head units and ratios you can reuse

Traditional figure drawing uses head units. The average adult is roughly seven to seven and a half heads tall. Fashion illustration pushes that to eight or nine. Real people cluster between six and a half and eight heads. If you want a character to read as unusually tall, you do not write "very tall." You write something like "seven point eight head units, narrow shoulders, long lower legs, small head relative to torso."

Add the relational markers that survive clothing changes and camera angles:

  • Shoulder width: roughly two head units for a broad frame, closer to one and a half for a lean frame
  • Arm span: approximately equal to total height in most adults
  • Elbow position: at or just below waist level
  • Wrist position: at the hip crease when arms hang naturally
  • Knee: roughly halfway between hip and floor
  • Hand length: about one head unit

These ratios hold up across wardrobe changes, which is exactly what you need when a character wears four outfits across twelve shots.

Real-world units for scene anchors

Abstract ratios only help if the set has references too. Keep a short list of standard dimensions and put two or three of them into your prompts or your reference images:

  • Doorway opening: about 2.0 to 2.1 m (80 in) tall, 0.8 m wide
  • Standard ceiling: 2.4 to 2.7 m
  • Counter and table height: 0.9 m
  • Chair seat: 0.45 m; bar stool seat: 0.75 m
  • Mirror lower edge in a dressing room: 0.9 to 1.0 m from the floor
  • Clothing rail: 1.7 to 1.8 m high, garments hanging to about 1.1 m

Once these anchors exist in your image, a viewer's brain instantly calibrates the human in the frame. Remove them and the same human looks like a scale model in a dollhouse.

Build a one-page proportion sheet

Create a single document for each recurring character containing: name, height in centimetres and feet, head-unit count, shoulder width, inseam, apparent shoe size, posture notes, and three reference images from different angles. This sheet is your single source of truth. Every prompt, every reference image, and every continuity check refers back to it.

Locking a Character Across Shots

Scale accuracy is only half the battle. The other half is identity consistency, which is where most AI video projects fall apart between shot three and shot nine.

Anchor your keyframes deliberately

Generate one clean, front-facing, neutral-light, mid-body framing of your character before anything else. That image is the anchor. Then create three derived anchors: a three-quarter view, a profile, and a full-body shot with the set's scale references visible. Lock these four images and reuse them as image-to-video seeds or as structural references for every subsequent shot.

When you generate new angles, always reference the nearest anchor rather than generating fresh from text. Text-only generation is where drift creeps in: the jaw widens, the shoulders narrow, the head grows. Chaining from a locked image keeps those ratios stable.

Wardrobe, hair, and continuity notes

Write continuity notes like a script supervisor would. Not "wearing a blazer" but "charcoal single-breasted blazer, sleeves ending at wrist bone, collar unbuttoned, no visible logos." Note the state of the garment per scene: jacket on or off, sleeves rolled, tie loosened. Hair gets the same treatment: parted left, tucked behind right ear, two strands loose at the temple.

These details matter more in AI video than in traditional film because the model has no memory. Every shot is a fresh interpretation, and any ambiguity gets resolved differently each time.

When height legitimately changes

Sometimes height should change: heels versus flats, a character standing on a step, a forced perspective shot. Handle those as explicit, documented exceptions rather than accidents. Note the delta in centimetres in your continuity sheet, and shoot the shots in matched pairs so you can check them side by side in the edit.

The worst outcome is an unintentional change that only becomes visible when both shots are cut together. Reviewing matched pairs on a timeline before you commit to a final render catches almost all of these.

Set Design for AI Video: Dressing Rooms, Backstage Areas, and Tight Interiors

Small interiors are the hardest thing to generate convincingly, because every object in the frame is a scale cue. Get three of them wrong and the whole space reads as miniature.

Ergonomics of a functional small room

A realistic dressing room is not a walk-in closet with a chair. It has a working triangle: mirror and counter, wardrobe rail, and seating, all within two steps of each other. Typical dimensions run 2.5 by 3 m for a functional single-occupant room, or 3 by 4 m for one shared by a small team. Ceilings are often lower than residential, around 2.4 m, with lighting fixtures adding visual weight at the top of frame.

Plan for these functional zones in your prompt: a mirror wall with bulb lighting, a counter deep enough for tools at 0.5 m, a rail that holds garments without crowding, and a chair or bench that a person can actually sit on with legroom. When the seat is too close to the counter, the human in the frame will look compressed and oversized at the same time, which reads as a scale error even if the viewer cannot articulate why.

Clutter budgets

Every object you add is a scale cue, for better or worse. A useful rule: three large anchor objects, five to seven medium objects, and a scattering of small items. Anchors establish scale (chair, rail, mirror). Medium objects add character (garment bags, a steamer, a stack of scripts). Small items add texture (brushes, pins, a water bottle).

Beyond that, clutter becomes visual noise that the model interprets unpredictably. If you find yourself generating objects that melt into each other, cut the small-item category entirely for that shot and reintroduce it in close-ups where the framing is tight enough to control.

Lighting tight interiors

Small rooms with a single window generate harsh falloff and unpredictable shadows. Three practical approaches:

  1. Practical lighting: mirror bulbs, ceiling panels, and a warm floor lamp give the model multiple soft sources and hide detail inconsistencies.
  2. High-key flat light: reduce shadow complexity when the space is busy, so the viewer focuses on the human rather than the geometry.
  3. Directional single source: use this only in close framing, where the falloff lands on the face rather than on ambiguous background objects.

Match your light direction to your established key direction across shots. A character lit from the left in one shot and the right in the next reads as a different location, even if the set is identical.

A Repeatable Production Workflow

Here is the sequence that holds up on real projects, from a two-shot test to a sixty-shot sequence.

Step 1: Assemble a reference board

Collect or generate ten to twenty images covering character angles, wardrobe, set layout, lighting, and colour palette. Do not start generating motion until this board is stable. The board is cheap to change; rendered video is not.

Step 2: Write the character and set bibles

One page each. The character bible holds proportions, wardrobe states, and continuity notes. The set bible holds dimensions, anchor objects, light direction, and a simple floor plan sketch. Keep both open in a second window while you work.

Step 3: Build the shot list with a scale chart

For every shot, record: framing (wide, mid, close), camera height, character height in frame as a fraction of total frame height, and which scale anchors are visible. Camera height matters enormously. A camera at 1.6 m with the character at 1.85 m creates a very different relationship than a camera at 0.9 m looking up.

Step 4: Generate anchors, then derivatives

Produce your four anchor images per character first. Approve them. Only then generate shot variations as image-to-video or as controlled derivatives. Never let a text prompt introduce a new interpretation of the character mid-sequence.

Step 5: Run cheap passes before expensive ones

Generate at low resolution with short duration to check composition, limb counts, and scale relationships. Fix problems here. It is dramatically faster to regenerate twelve low-resolution tests than three high-quality renders.

Step 6: Assemble and audit continuity

Cut everything to a rough timeline before final rendering. Watch it once for identity, once for scale, and once for lighting direction. Most errors only become obvious in sequence, not in isolation.

Tooling That Helps Without Adding Chaos

You do not need a large stack. You need one image generator, one video model, one control layer, and one editor.

  • Image generation for anchors and reference boards, ideally with seed locking and image-to-image strength control.
  • Video generation that accepts an image as the first frame or as a structural guide. This is the single most important feature for consistency.
  • Pose and depth control using pose estimation, depth maps, or edge detection to force the model toward a specific silhouette and body scale rather than an approximate one.
  • Upscaling as a final pass, applied consistently so grain and detail do not shift between shots.
  • A standard non-linear editor for continuity auditing. Free editors do this perfectly well; the discipline matters more than the software.

Add tools only when a specific, repeatable failure demands it. Every additional model introduces another interpretation of your character.

Mistakes That Destroy Believability

  • Describing height with adjectives. "Very tall" is meaningless to a model. Use head units and centimetres.
  • Generating each shot from scratch. Start from an approved anchor image whenever the character or set repeats.
  • Forgetting the scale anchors. A dressing room with no door, no chair, and no counter has nothing to calibrate against.
  • Changing camera height unintentionally. Sudden low-angle shots make characters look taller and sets look larger.
  • Over-cluttering small rooms. Visual noise produces melted objects and inconsistent geometry.
  • Ignoring light direction. Inconsistent key light reads as a location change.
  • Rendering before editing. Rough cuts at low quality reveal continuity errors long before a full render does.
  • Skipping the wardrobe state log. Jacket on or off is a continuity fact, not a stylistic choice.

Ethics, Likeness, and Real People

There is a reason accurate proportions matter beyond craft. When a generated figure resembles a real, identifiable person, you are creating a likeness, and that carries legal and ethical weight regardless of intent. Many jurisdictions treat synthetic likeness as protected, and platforms increasingly require disclosure of synthetic media.

The practical path is straightforward: build fictional archetypes rather than recreating real individuals. Give your character a distinct proportion sheet, a specific wardrobe, and a name of your own. If a project genuinely requires depicting a real person, secure documented permission, keep the scope narrow, and follow the disclosure rules of every platform you publish to. Accuracy in scale is craft. Accuracy in likeness is a responsibility.

FAQ

How do I stop a character from changing height between shots?
Lock one full-body anchor image that includes a scale reference such as a doorframe, and derive all other shots from it. Record camera height and character height as a fraction of frame in your shot list, and check matched shots side by side in an edit before rendering.

What is a head unit and why does it matter?
A head unit is the height of the character's head used as a measuring module. Adults are roughly seven to seven and a half heads tall. Specifying head-unit counts gives a model a ratio to preserve rather than an adjective to interpret.

How large should a dressing room set be?
Around 2.5 by 3 m for one occupant, up to 3 by 4 m for a small team, with a ceiling near 2.4 m. The key is not the exact figure but the relationships: counter at 0.9 m, rail at 1.7 to 1.8 m, seating with real legroom.

Do I need 3D software for this?
No. A floor plan sketch, a proportion sheet, and consistent scale anchors in your reference images get you most of the way. 3D tools help when a camera move needs to be exact, but they are not a prerequisite.

How many reference images should a character have?
Four anchors minimum: front, three-quarter, profile, and full body with scale references. Add one per wardrobe state if the outfit changes significantly across the sequence.

What is the fastest way to catch scale errors?
Cut a rough timeline at low resolution and watch it three times: once for identity, once for scale, and once for lighting. Errors that are invisible in a single frame become obvious in sequence.

Alexander

Alexander