Why structure beats vocabulary in AI video prompting
Ask ten people to describe the same shot and you will get ten prompts that produce ten different clips. That is not a failure of the models. It is a failure of the brief. A generative video engine is an ambiguity resolver: wherever your instruction is precise, it follows you, and wherever it is vague, it quietly substitutes the statistical average of everything it has seen. A generic face, a generic street, a generic slow push-in.
The practical consequence is that prompting skill is not about writing beautifully. It is about closing the distance between the image in your head and the default behaviour of the system you are feeding. Four principles carry most of the weight:
- Specificity over intensity. "Low-angle shot, warm backlight rimming the shoulders, shadow detail retained" outperforms "incredibly beautiful cinema lighting."
- Structure over length. A sixty-word prompt organised into named slots usually beats a two-hundred-word paragraph of impressions.
- Constraints over wishes. Telling the model what to avoid — readable signage, lens flares, handheld shake — is often stronger than adding another adjective.
- Iteration over perfection. The first render is a hypothesis, not a verdict. Change one variable, render again, compare.
Everything else in this guide is an elaboration of those four ideas.
The six building blocks of a workable video prompt
A prompt that survives contact with a real project has six representable slots. They can be one sentence each or compressed into three tight sentences, but each one should be answerable by reading the prompt back.
1. Subject and action
Say who or what, doing what, in what emotional register. Put this first, because most text encoders weight early tokens more heavily and because composition follows subject. Be concrete about age range, build, wardrobe and posture instead of stacking superlatives. "A locksmith in his forties, denim apron, crouching to work a lock" gives the model something to draw on; "a cool guy" gives it nothing.
One primary action per clip. If the shot needs three actions, it needs three shots.
2. Camera and lens
Camera language is the highest-leverage vocabulary you have, because it controls both composition and perceived production value. Useful dimensions:
- Shot size: extreme close-up, close-up, medium, wide, establishing.
- Angle: eye level, low angle, high angle, tilted frame, over-the-shoulder.
- Movement: locked-off, slow push-in, pull-back, pan, tracking, crane up, orbit, handheld follow.
- Lens character: 24mm wide, 50mm normal, 85mm portrait, macro, anamorphic with mild horizontal flare.
Combine one shot size, one angle and at most one movement. Prompts that request a push-in, a pan and an orbit simultaneously produce mush, because the model has to average three incompatible camera paths.
3. Light, palette and grade
Lighting decides whether output reads as amateur or cinematic more reliably than any other single factor. Specify direction, quality and colour:
- Direction: backlit, side-lit, top-lit, frontal soft.
- Quality: hard sun, soft overcast, diffused window light, practical neon.
- Palette: warm amber interior, cool desaturated exterior, teal-and-orange, monochrome with one accent.
- Grade language: "muted highlights, lifted blacks, slight green cast in the shadows" is actionable; "moody" is not.
4. Motion and timing
Describe camera motion and subject motion separately, then add pace. Words such as slow, deliberate, frantic, halting, at half speed reshape temporal sampling. For short clips, one continuous motion reads far better than an implied cut. If the engine exposes duration, match it to the action: three seconds for a gesture, eight for a walk across frame.
5. Style anchor
Style anchors are shorthand: grainy 16mm documentary, hand-painted cel animation, high-gloss automotive commercial. The trap is over-anchoring. Three or four references in one prompt fight each other and the output lands somewhere incoherent. Pick one primary anchor and let a single secondary detail support it.
6. Constraints
A short exclusion list catches recurring defects: extra fingers, warped faces, jittery motion, blurred text, watermark artefacts, duplicated limbs, frame flicker. Keep it to five to eight targeted items. Long exclusion lists sometimes suppress legitimate content along with the defect.
From slots to shots: weak prompts versus strong prompts
Here is the same idea written two ways. Weak:
A cool cyberpunk scene, woman walking, very cinematic, amazing lighting, 4K, masterpiece.
Working:
Medium tracking shot at eye level, 35mm, following a woman in a translucent raincoat as she walks left to right through a narrow alley. Neon signage reflects in puddles; magenta side light from the left, cyan fill from the right; light drizzle. Slow, deliberate pace, coat billowing slightly. Cinematic night grade, deep blacks, highlight detail retained, shallow depth of field. No text, no lens flare, no camera shake.
The second version fixes four things the first leaves open: composition, direction of travel, the number and position of light sources, and pacing.
Ordering discipline
Keep the slot order stable: subject, camera, light, motion, style, constraints. A stable order makes prompts debuggable. When a render fails, you can usually point at the slot that caused it rather than rewriting everything at once.
Two-pass writing
Write the prompt as a numbered block for your own reference — six lines, one per slot — then flatten it into a paragraph before pasting it into the tool. The numbered pass forces you to notice an empty slot. The flattened pass keeps the text natural for the encoder.
Matching prompt style to the engine you are using
Engines differ in how much they reward verbosity, how literally they interpret camera directives and how strongly they enforce physical plausibility. Treat the following as habits, and verify each one with your own test renders.
Photoreal engines
Photographic language works best: focal length, aperture feel, film stock, lighting ratios. These models often respond better to restraint on style adjectives, because heavy stylisation pushes output toward illustration. Lean on texture descriptors — skin pores, fabric weave, condensation, dust in air. If faces drift or deform, add a framing directive: "medium close-up, subject centred, face fully visible, shoulders in frame."
Motion-heavy engines
Systems built for dynamic movement reward explicit motion verbs and a clear statement of subject scale in frame. Describe the path: "enters from frame left, moves diagonally toward camera, exits right." Avoid stacking simultaneous camera moves. If output jitters, lower the requested speed and simplify the background before you touch anything else.
Stylised and animation engines
Animation-oriented systems reward style tokens more generously than photoreal engines do: line weight, shading model, colour script, silhouette clarity. "Two-tone cel shading with hard shadow edges and flat background colour" is far more useful than "anime style." Keep backgrounds simpler, because stylised models have less tolerance for dense detail and will render complexity as visual noise.
Image-to-video and keyframe workflows
When you start from a still frame, the prompt's job changes: you are no longer describing a scene, you are directing a performance. Describe motion and camera only. Repeating wardrobe, lighting and location details often makes the model reinvent the frame you already liked. In a first-and-last-frame workflow, describe the transition — "hand starts open and closes into a fist" — and let the two images hold identity.
Camera and lighting vocabulary that reliably changes output
A personal vocabulary beats a giant list. Build a short list of terms you have tested and can trust.
Camera terms worth keeping: locked-off tripod, slow push-in, dolly out, shoulder-mounted follow, crane rise, parallax foreground, deep focus, shallow focus, 35mm normal perspective, 85mm compression.
Lighting terms worth keeping: key from frame left, soft fill, hard rim, practical sources in frame, bounced ceiling light, overcast diffusion, golden-hour backlight, night interior with motivated lamp.
Grade terms worth keeping: lifted blacks, crushed shadows, halation on highlights, desaturated midtones, warm highlight roll-off, cool shadow cast, subtle grain.
The trick is that each of these maps to an observable change in the output. If a term does not change anything when you swap it, it is decoration. Remove it and free up attention for a term that does.
Keeping characters and locations consistent across shots
Consistency is where casual experiments and professional sequences diverge. Across a multi-shot scene, faces, wardrobe, props and locations must read as the same world.
Character blocks
Write a reusable character block: age range, build, hair colour and style, facial hair, distinguishing marks, default wardrobe, silhouette shape. Paste that block verbatim into every prompt featuring the character. Do not paraphrase, and do not "improve" the wording between shots. Small wording changes move the output in ways that are easy to see and hard to undo.
Location and prop blocks
The same discipline applies to sets. Define each location once — "converted warehouse loft, exposed brick, tall industrial windows on the left, scuffed concrete floor" — and reuse the exact wording. Props that appear in several shots, such as a specific phone, bag or car, should be described once with colour, shape and condition, then referenced identically.
Frame and grade blocks
Aspect ratio, frame rate feel, grain level and colour grade should also be fixed in a block. A sequence where three shots are vertical and two are wide, or where two shots are graded cool and three warm, reads as a mistake rather than a style choice unless the change is motivated by story.
A continuity checklist
Before rendering a sequence, confirm that:
- Every character block is identical across prompts.
- Location wording is unchanged between shots in the same scene.
- Lighting direction is consistent unless the story calls for a change.
- Time of day matches across adjacent shots.
- Aspect ratio, grain and grade match across the whole sequence.
The iteration loop: change one variable at a time
Prompting is a loop, not a single act. The loop only works if you isolate variables.
One variable per render
If you change camera, lighting and style at the same time, you learn nothing about which change fixed the shot. Identify the slot that is failing, adjust only that slot, re-render and compare side by side. This feels slower for the first ten clips and dramatically faster for the next hundred.
Seeds, duration and aspect ratio
Locking the random seed lets you separate prompt effects from sampling noise. Duration and aspect ratio are structural decisions, not cosmetic ones: a 9:16 vertical frame changes what the camera can physically show, so choose the delivery format before you tune the prompt. A wide establishing shot composed for a horizontal frame will not survive a vertical crop, no matter how many adjectives you add.
Take logs
Keep a plain-text log with prompt, engine, settings, seed and a one-line verdict. A simple format works:
take 014 | locked seed | close-up, hard rim light | face stable, rim too hot
take 015 | locked seed | close-up, softer rim | good, keep
After twenty clips you will have a personal reference sheet worth more than any generic prompt collection, because it describes the specific systems you actually use.
Common mistakes and how to fix them
Adjective stacking. Ten praise words dilute each other until none of them steer the output. Fix: keep two or three sensory descriptors and spend the rest of the prompt on technical specifics.
Contradictory camera moves. A push-in plus a pan plus an orbit. Fix: one movement per clip, chosen to serve the beat.
Style soup. Four references pulled from four different visual traditions. Fix: one primary anchor plus at most one supporting detail.
Ignoring the delivery format. Composing a wide shot for a vertical feed, or a busy wide shot for a small phone screen. Fix: decide aspect ratio first, block the frame second, write the prompt third.
No exclusion list. Recurring artefacts that you keep editing around instead of prompting away. Fix: five to eight targeted exclusions, reviewed every few weeks.
Expecting features an engine does not have. Reliable lip-sync dialogue, precise on-screen text and long unbroken takes are still uneven across systems. Fix: design around the strengths you have observed and solve the rest in the edit.
Rewriting everything after one bad take. Fix: single-variable changes and a take log.
Chasing a fixed idea of the perfect clip. Fix: judge renders against the shot brief, not against taste.
A repeatable workflow from script to finished sequence
- Break the script into shots. One action and one camera idea per shot.
- Write a two-line shot brief. Plain language: what must the audience see and feel here?
- Fill the six slots. Subject, camera, light, motion, style, constraints.
- Flatten and render three takes with different seeds, keeping everything else identical.
- Evaluate against the brief, not against your mood. A clip that communicates the beat and cuts cleanly is a pass even if it is not beautiful in isolation.
- Adjust exactly one slot and re-render.
- Lock character, location and grade blocks. Reuse them verbatim for the rest of the sequence.
- Edit for rhythm. Cut on motion rather than on frame boundaries; trim the first and last frames of most generations.
- Archive the prompt with the final clip. Six months later, the prompt is the only documentation that matters.
FAQ
How long should a video prompt be? Usually forty to ninety words. Long enough to cover the six slots, short enough that no directive fights another.
Should I write prompts in my own language or in English? Use whichever language the engine handles best. Most major systems perform strongly in English, and translation can shift nuance in lighting and camera terms. Keep your own prompt library in a single language so your notes stay comparable.
Can I name a well-known director or artist as a style reference? It is unreliable and ethically awkward. Describe the visual traits instead: palette, grain, line quality, lighting ratio, camera height. Traits transfer between engines; names do not.
Why does the same prompt give different results every time? Sampling randomness, plus any seed variation. Lock the seed while you are tuning and vary it deliberately when you are exploring.
How do I get stable faces across shots? Verbatim character blocks, consistent lighting direction, and a reference image or first-frame input when the engine supports it.
What do I do when the model ignores a slot entirely? Move that slot earlier in the prompt, express it in physical terms rather than stylistic ones, and remove any competing instruction. If it still fails, the engine likely cannot do it — solve it in the edit instead.
When should I stop iterating? When the clip satisfies the shot brief and cuts cleanly with its neighbours. Further passes on a single clip rarely improve the finished sequence, and they consume the time you need for the shots that actually matter.
Where to take this next
Prompting improves fastest when you treat it as an engineering habit rather than a creative mood. Build your own six-slot template. Keep a take log so your experiments compound instead of evaporating. Lock your vocabulary so that a working phrase stays working. And review your render queue weekly to spot the artefacts you keep fixing by hand — those are the ones that belong in your exclusion list.
The practitioners who get predictable results are not writing more words. They are writing more structured ones, they test one variable at a time, and they refuse to paraphrase a phrase that already works.


