Why prompt craft still beats model shopping
Every few weeks a new image or video model arrives, and the temptation is to assume the upgrade will fix everything. It rarely does. The gap between a mediocre generation and a shot you would actually put in a client edit is almost always a prompt gap, not a model gap. Two people can use the same model and the same seed and get wildly different results because one of them described a subject and the other described a moment.
A useful mental shift is to stop thinking of prompts as search queries and start thinking of them as shot briefs. A search query retrieves something that already exists. A shot brief tells a crew what to build: who is in frame, what they are doing, where they are standing, what the light is doing, what lens is on the camera, and how the camera moves. When your prompt contains those layers, the model has enough constraints to produce something coherent instead of something merely plausible.
The other reason prompt craft matters is reproducibility. Freelancers, small studios, and solo creators all run into the same problem: a client approves a look, and three weeks later nobody can recreate it. If your prompt is a single sentence of vibes, that look is gone forever. If your prompt is a structured template with named variables, you can regenerate the look on a different model, at a different aspect ratio, or in a different scene.
This guide walks through that structure, with concrete examples you can adapt. It applies equally to still image generation and to text-to-video, with notes where the two diverge.
The anatomy of a cinematic prompt
Most strong prompts can be broken into six layers. You do not need all six every time, but knowing which layer is missing is the fastest way to diagnose a weak result.
1. Subject and wardrobe
Be specific in a way that is visual, not narrative. "A determined detective" tells the model almost nothing. "A 40-year-old detective, weathered face, three-day stubble, charcoal wool coat with a frayed collar, silver watch on the left wrist" gives it something to render. Wardrobe details do double duty: they anchor identity and they give the model texture to work with, which is where image quality actually lives.
Avoid stacking contradictory identity traits. "Young" plus "deep wrinkles" plus "childlike" produces mush. Pick three to five identity anchors and hold them constant across every shot of the same character.
2. Action and verb choice
Verbs shape pose more than adjectives do. "Standing" is neutral and often produces a stiff mannequin. "Mid-stride, left foot forward, coat hem lifting" produces weight and motion. In video, the verb also suggests the timing: "slowly turning her head toward the window" implies a duration, while "snaps her head toward the window" implies a fast cut.
One strong verb beats four vague ones. If you need two beats of action in a single clip, expect the model to blend them, and plan for that rather than fighting it.
3. Environment and depth cues
Environment is where most prompts collapse. "In a city" gives you a generic skyline. "On a rain-slicked side street at night, neon signage reflected in standing water, fire escape receding into fog above" gives you depth layers: foreground reflection, midground street, background architecture. Models respond well to explicit foreground, midground, and background cues because it forces them to construct spatial depth rather than paint a flat backdrop.
4. Lighting and color temperature
Lighting is the single highest-leverage layer. Name the source, the direction, and the quality:
- Source: practical lamp, overcast sky, single window, neon sign, firelight
- Direction: backlit, side-lit from camera left, top-down, rim light from behind
- Quality: hard shadows, soft diffusion, volumetric haze
Then add temperature. "Cool blue moonlight on the left, warm amber practical on the right" gives you a two-tone scheme, which reads as professional immediately because it is how real sets are lit.
5. Lens, format, and grade
Camera language is a shortcut to style. "85mm portrait lens, shallow depth of field, slight background compression" reads as a considered photograph. "Wide 24mm, deep focus, slight barrel distortion" reads as documentary. Add a film stock or grade reference if you want a consistent look across a project: muted teal shadows, warm highlights, fine grain.
Be careful with brand names. Some models handle them well, others produce an approximation that looks like a knockoff. Describing the look — contrast curve, halation, grain structure — is usually more reliable than naming a product.
6. Motion clause (video only)
For video, append a short clause describing camera behavior and subject motion separately. Conflating them is a common error: "the camera slowly pushes in while the dancer spins" gives the model two simultaneous instructions, and results improve when you phrase them as separate beats with clear priorities.
A reusable prompt template and worked examples
Here is a template that has survived contact with several different models. Bracketed items are variables.
[SUBJECT with 3-5 identity anchors], [WARDROBE], [SINGLE ACTION VERB PHRASE],
[ENVIRONMENT with foreground / midground / background],
[LIGHTING source + direction + quality], [COLOR TEMPERATURE],
[LENS + DEPTH OF FIELD + FORMAT], [GRADE or FILM LOOK], [MOOD WORD]
A filled example for a still:
A 34-year-old botanist, freckles across the nose, dark curly hair tied back,
linen shirt with rolled sleeves, leaning over a specimen tray, turning a leaf
with tweezers, inside a glass greenhouse, foreground fern fronds out of focus,
midground steel worktable, background misted glass panels receding,
soft diffuse daylight from above, warm green-yellow casts, 50mm lens, medium
depth of field, 3:2 frame, muted natural grade, quiet concentration
And the same character in a video prompt:
A 34-year-old botanist, freckles across the nose, dark curly hair tied back,
linen shirt with rolled sleeves, carefully lifting a specimen jar and holding it
to the light, inside a dim greenhouse at dusk, foreground hanging vines,
midground worktable with scattered notes, background warm lamp glow,
practical tungsten lamp from camera right, cool dusk ambient from the left,
35mm lens, shallow focus, slow handheld drift to the right, subject stays
centered, gentle grain, contemplative
Notice the differences. The video version drops the most granular wardrobe texture (motion and detail compete for the model's attention), simplifies the lighting to two clear sources, and adds a camera clause with an explicit "subject stays centered" anchor to reduce drift.
For a wider establishing shot, invert the emphasis: environment and atmosphere first, subject small in frame, and a slower camera move. Establishing shots tolerate far more background complexity than character shots, so this is the place to spend your detail budget.
Negative prompts, weighting, and controlled accidents
Negative prompts are most useful when they target a specific failure you have already observed, not when they are a laundry list copied from a forum. If your previous render gave the subject six fingers, add a targeted negative. If it looked fine, adding twenty negatives only risks removing things you wanted.
A practical short negative list for character work:
extra limbs, fused fingers, warped face, duplicate head, text, watermark,
plastic skin, oversharpened, blown highlights
For environments, the failure modes differ, so the list should too:
flat horizon, empty foreground, cluttered midground, repeating textures,
visible tiling, muddy shadows, washed-out sky
Weighting is the other half of control. Most interfaces let you emphasize terms, either with syntax or by reordering. Weight the layer that the shot depends on: if identity is critical, weight the wardrobe and face anchors. If atmosphere is critical, weight the lighting and environment. Do not weight everything — uniform emphasis is the same as no emphasis.
One underused technique is deliberate imperfection. Adding a small amount of grit — "slight lens flare across the lower third," "dust motes visible in the beam," "subtle motion blur on the trailing hand" — makes generations read as captured rather than computed. Perfect symmetry and clean gradients are the visual signature of synthetic output; small asymmetries fix that faster than any style keyword.
Keeping characters recognizable across a sequence
Consistency is the hardest problem in AI video, and no single trick solves it. A layered approach works better than any one method.
Build a character bible. Write down your three to five identity anchors in a fixed order and never reorder them. Models are sensitive to token position, so a stable phrase order reduces identity drift across shots.
Generate a reference sheet first. Produce a still with neutral lighting, the character facing camera, and the full wardrobe visible. Use it as an image reference or as the first keyframe for subsequent generations. Text alone is a weak identity signal; text plus a reference image is dramatically stronger.
Separate identity from performance. Keep the identity block identical in every prompt and change only the action, environment, and camera clauses. If identity drifts, you know exactly which part of the prompt to blame.
Accept controlled variation. A character can look slightly different across shots if the lighting and angle differ, and audiences generally accept this. What breaks the illusion is a change in the anchors themselves. If the coat changes color between two shots in the same scene, that is a continuity error; a slightly different jawline under different lighting is not.
Test at the target aspect ratio. Identity behavior often changes between square and widescreen framing because composition rules push the model toward different crops. Lock your aspect ratio early.
Scene continuity and environment locking
Environmental continuity fails in three predictable ways: the props change, the time of day shifts, and the spatial relationships between objects flip.
To prevent prop drift, name the three to five objects that define the space and reuse that exact list in every prompt for the scene. "Steel worktable, brass microscope, stack of leather-bound notebooks, single hanging bulb" is a describable set. Repeating it verbatim does most of the work.
Time of day is best communicated through light rather than the word "evening." "Deep blue ambient with warm lamp pools" is more actionable than "dusk," and it also tells you exactly what to keep constant.
Spatial relationships need explicit anchor phrases: "worktable on the left, door frame on the right, window behind the subject." When you change camera angle between shots in the same scene, keep the anchors and change only the viewpoint clause, for example "viewed from the doorway" or "low angle from table height." This is the text equivalent of a shot list.
Directing camera motion and performance with text
Video models respond to camera language surprisingly well when the instruction is unambiguous and singular. A short vocabulary covers most needs:
- Push in / pull out — changes intensity. Best for emotional beats.
- Pan left or right — reveals environment. Best for establishing shots.
- Tilt up or down — reveals scale. Best for architecture and crowds.
- Tracking alongside — keeps subject at constant size while environment moves.
- Arc around subject — creates dimension. Works best with a clear single subject.
- Handheld drift — adds realism and hides minor artifacts.
Rules that consistently improve results:
- One primary move per clip. Two moves in four seconds reads as a mistake, not a style.
- Specify speed. "Slow" and "fast" produce different amounts of temporal smear; unqualified moves default to medium, which often looks synthetic.
- Anchor the subject. Phrases like "subject stays centered" or "subject remains in the lower right third" reduce unwanted drift.
- Separate subject motion from camera motion so the model does not average them.
Performance is conveyable too, but through physical detail rather than emotion words. "Shoulders drop slightly, eyes flick to the left, jaw tightens" is actionable. "Sad and conflicted" is not. If you need an emotional read, describe the body and let the viewer supply the feeling.
Mixing styles and adapting prompts across models
Style blending is easiest when you name the visual traits instead of the sources. "High-contrast graphic lighting with limited palette and heavy shadow shapes" survives a model swap. "In the style of a specific graphic novel artist" may work brilliantly on one model and produce a muddy pastiche on another.
When moving a prompt between models, change three things first:
- Length. Some models reward dense prompts, others start ignoring later clauses. If results get worse as you add detail, cut the last third.
- Negative prompt handling. Some models treat negatives literally, others barely register them. Retest your core negatives after any model change.
- Stylistic burden. A model with strong built-in realism needs less film-language; a highly stylized model needs more scene description to stay grounded.
Keep a personal prompt library organized by scene type rather than by model. "Character close-up, warm interior," "exterior establishing, overcast," and "action beat, high shutter" are reusable across tools; model names are not.
Common mistakes, fixes, and a troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
| Stiff, mannequin-like subjects | No action verb; neutral pose language | Replace "standing" with a mid-motion phrase |
| Flat, poster-like images | No foreground/midground/background separation | Add three explicit depth layers |
| Identity drifts between shots | Identity block reordered or paraphrased | Freeze the identity phrase verbatim |
| Warped hands and faces | Too much competing detail; hidden subject | Move subject larger in frame, cut secondary descriptions |
| Muddy lighting | Multiple unspecified light sources | Name two sources max with direction and quality |
| Motion looks synthetic | Two simultaneous camera moves | Reduce to one primary move, add speed and anchor |
| Style looks like a knockoff | Brand or artist name used as a shortcut | Describe contrast, palette, grain, and shadow shape |
| Prompt stops working after edits | Later clauses diluting earlier ones | Reorder so the most important layers come first |
A fast diagnostic routine: generate four variations with only one layer changed. If nothing changes, that layer is probably not being read. If everything changes, your layer is too broad. This takes five minutes and saves hours of blind tweaking.
FAQ: prompt writing for AI art and video
How long should a prompt be? Long enough to cover the six layers, short enough that each clause does work. For character shots, 60–90 words is a reliable range. Establishing shots tolerate 100–140. If adding words stops changing the output, you have passed the useful limit.
Do I need a negative prompt at all? Only if you have a specific failure to suppress. A blanket negative list is a blunt instrument that occasionally removes desirable texture.
Why does the same prompt give different results tomorrow? Models get updated, and many interfaces introduce subtle randomness in the sampler. Save your seeds when you get a keeper, and note the interface version if your tool exposes it.
Should I write prompts in my own language? Write in the language your model handles best for the specific vocabulary you need. Lighting and lens terms are often strongest in English, but scene and dialogue context may be better in your native language. Mixed-language prompts are acceptable in most interfaces and often outperform awkward translations.
How do I get consistent characters without reference images? You can get close with a frozen identity block and matching lighting, but expect some drift. Reference images or a fixed keyframe are far more reliable for any sequence longer than two shots.
What is the fastest way to learn a new model? Run the same three prompts on it: one character close-up, one wide establishing shot, and one fast action beat. Those three expose most of a model's biases within ten minutes.
Can one prompt serve both image and video? Usually with edits. Strip granular texture detail, simplify lighting to two sources, and add a single camera clause with a speed and an anchor. The identity and environment blocks can stay identical, which is what keeps a sequence coherent.
A repeatable production loop
Prompting gets fast once you stop improvising. The loop that works: define the six layers in a text file, generate a reference still, lock identity and environment blocks, iterate on one layer at a time, then build shots by changing only the action and camera clauses.
Treated this way, a prompt is less a magic phrase and more a piece of production documentation — shareable, versionable, and reusable across projects and tools. That is what turns occasional good generations into a reliable pipeline, and it is the difference between hoping a model cooperates and directing it.

