AI video generation has moved well past the novelty stage. Typing a single poetic sentence and getting a moving image back is no longer impressive on its own. What separates a clip that gets scrolled past from one that holds attention is rarely the model you chose — it is the thinking that happened before you typed anything at all.
Prompt engineering for video is not the same discipline as prompting a text model. A language model wants context, nuance, and instruction. A video model wants a compressed visual specification: who is in frame, what they are doing, where the camera is, how the light behaves, and what must stay identical from the previous shot. Those are different muscles, and most creators only train one of them.
This guide is a practical system for building video scripts that survive generation. It covers prompt anatomy, story spines, camera language, continuity, pacing, a full worked example, and the mistakes that quietly ruin otherwise good projects.
The Anatomy of a Scene Prompt
A scene prompt is a compressed shot description, not a paragraph of prose. When it works, a reader could storyboard from it without asking a single clarifying question.
The five building blocks
Almost every reliable prompt can be decomposed into five parts:
- Subject — the specific who or what, with the two or three visual details that matter most. "A weathered lighthouse keeper in a wool coat" beats "an old man."
- Action — a single visible verb happening in the frame. Not a backstory, not an intention, an observable motion.
- Environment — location, time of day, weather, and the spatial relationship between subject and surroundings.
- Camera — shot size, angle, movement, and lens feel. This is the block most creators skip, and it is the block that most reliably changes output quality.
- Look — lighting, palette, texture, film stock or render style, and the emotional temperature.
Two extra layers are optional but powerful: continuity anchors (the details that must match neighbouring shots) and negative constraints (what should not appear).
Freeform prompts versus structured prompts
Freeform prompting is how you explore. You throw six variations at a model in ten minutes and see which direction feels alive. That stage is genuinely creative and should not be over-engineered.
Structured prompting is how you produce. Once a direction is chosen, lock the structure and stop improvising. A workable production template looks like this:
[SHOT TYPE] subject + defining details,
[ACTION] present-tense visible verb,
[ENVIRONMENT] location, time, weather, spatial notes,
[CAMERA] angle, movement, lens, speed,
[LIGHT & LOOK] source, quality, palette, texture,
[CONTINUITY] must-match details from prior shots,
[AVOID] unwanted artefacts, styles, or elements
Keeping the same field order across a whole project does two things: it makes your own review faster, and it produces more stable results because the model receives the same information in the same sequence every time.
Build the Story Spine Before You Write a Single Prompt
The most common failure in AI video work is generating beautiful shots that do not add up to a story. The fix is unglamorous: outline first, prompt second.
Logline, beats, shot list
Start with one sentence. "A stranded astronaut repairs a beacon to signal a rescue that may already be too late." That sentence is your filter — if a shot idea does not serve it, it does not belong in the project.
Next, break the logline into four to seven beats. Beats are emotional or informational turns, not shots. A beat might be "hope fades as the beacon fails a second time." A shot is "close-up on a hand tightening a connector, sparks dying."
Finally, convert beats into a shot list. Each shot gets one line in the structured template above, plus an estimated duration. Only now do you start generating.
How many shots do you actually need
A useful rule: for a sixty-second piece, eight to fourteen shots. Fewer than eight and the edit feels sluggish; more than fourteen and individual shots are too short to register as images rather than flickers.
Resist the temptation to generate fifty options per shot. Generate two or three, pick one, move on. Projects die in the selection phase far more often than they die in the generation phase.
Directing With Words: Camera Language That Models Understand
Camera vocabulary is the highest-leverage addition to a prompt because it controls framing and energy without adding content the model has to invent.
Shot size, angle, and movement
Use standard, well-documented terms and combine at most two movement ideas per shot:
- Shot size: extreme wide, wide, medium wide, medium, medium close-up, close-up, extreme close-up, insert.
- Angle: eye level, low angle, high angle, overhead, over-the-shoulder, Dutch tilt.
- Movement: static lock-off, slow push in, pull back, lateral tracking, handheld follow, crane rise, orbit, whip pan.
"Medium close-up, slow push in, eye level" is dramatically better than "dramatic shot of her looking worried." The first tells the model what to draw. The second asks it to guess what drama means.
Lens, light, and palette
Lens and lighting language carries an enormous amount of meaning per word:
- Lens feel: 24mm wide with visible distortion, 50mm neutral, 85mm compressed portrait, macro detail, anamorphic flare.
- Light source: window light, overcast daylight, firelight, sodium street lamps, single practical lamp, hard sun with deep shadows.
- Quality: soft and diffused, harsh and directional, low-key with deep shadow, high-key and even.
- Palette: desaturated teal and grey, warm amber and dusty orange, cold blue night with magenta accents.
Pick a palette per project and reuse it in every prompt. Consistent palette is one of the cheapest ways to make unrelated shots feel like one film.
What models tend to ignore
In practice, video models drop or distort:
- Text — signage and readable letters are unreliable; avoid shots that depend on them.
- Hands and small objects — extreme close-ups of complex manual action often warp. Frame wider or hide the action behind motion.
- Crowds — background extras melt into each other when the camera moves.
- Multiple simultaneous actions — one clear action per shot is far more stable than three.
- Precise spatial logic — "the cup to the left of the knife behind the bowl" is usually lost. Simplify or reframe.
Knowing these limits changes the script itself. Write around the weaknesses rather than fighting them in post.
Consistency Across Shots: The Hardest Problem in AI Video
Generating one gorgeous shot is easy. Generating eight that look like they belong to the same film is where most projects collapse.
Build a character and world bible
Write a short document with locked descriptions:
- Character: age range, build, hair, distinguishing features, wardrobe with specific colours and materials, one signature prop.
- World: architecture style, era, dominant materials, weather pattern, recurring colour, one visual motif.
Then reuse those exact phrases, word for word, in every prompt where the character or location appears. Paraphrasing is the enemy of consistency. If the bible says "faded olive canvas jacket with a torn left cuff," do not later write "green jacket" and expect the model to connect them.
Use references and seeds deliberately
Where a model supports image references or seed locking, treat them as production assets rather than random luck:
- Generate a hero image for each character and location first, then use it as the reference for all subsequent shots.
- Lock a single seed for a sequence when you need near-identical camera behaviour.
- Keep a reference folder with one approved still per character, per location, per key prop.
When a shot drifts, compare it against the reference and ask which specific element changed — jawline, wardrobe colour, lighting direction. Fix one variable at a time. Changing three prompt fields at once makes diagnosis impossible.
Switching models mid-project
Mixing tools is normal, especially when one model handles faces better and another handles landscapes better. The cost is stylistic drift. Manage it by:
- Assigning models by shot function, not by mood. Model A for character close-ups, Model B for wide environments.
- Normalising in post with a shared colour grade, grain layer, and aspect ratio.
- Re-testing the character bible on any new model before committing a sequence to it.
A consistent grade will hide far more model drift than most creators expect.
Pacing and Emotion Without a Composer or an Editor's Timeline
Editing controls rhythm, but your shot list already decides most of it. Build pacing into the prompt plan.
Rhythm is shot length
A practical pattern for a sixty-second narrative piece:
- 0–5s: one wide establishing shot, static or slow push. Set place and tone.
- 5–20s: four to six medium and close shots, two to four seconds each. Establish the character and the problem.
- 20–40s: the complication. Shorten shots to one to two seconds, increase movement, tighten framing.
- 40–55s: the turn. One longer shot, four to seven seconds, often a slow push or a hold. Let it breathe.
- 55–60s: resolution image. One shot, often a pull back or a cut to black.
Notice that you are writing shot durations in the script, not discovering them in the editor. That is the difference between an AI video that feels authored and one that feels assembled.
Emotional verbs and micro-actions
Abstract emotion prompts ("sad," "tense," "hopeful") do very little. Embodied micro-actions do a lot:
- Instead of sad: "she exhales slowly, shoulders dropping, gaze drifting to the floor."
- Instead of tense: "his jaw tightens, fingers pressing flat against the table edge."
- Instead of hopeful: "a small breath of relief, eyes lifting toward the light."
Pair these with light and camera. A slow push combined with a single downward glance reads as grief. The same glance with a hard pull back and cold light reads as defeat. The performance is the body; the meaning is the camera.
A Worked Example: Sixty-Second Beacon Teaser
Let us walk the astronaut logline through the full pipeline.
Logline: A stranded astronaut repairs a beacon to signal a rescue that may already be too late.
Beats:
- Isolation established.
- The beacon is dead.
- Repair attempt succeeds briefly.
- It fails again — hope collapses.
- She tries once more anyway.
- Distant light answers, ambiguous.
Shot list excerpt:
| Shot | Duration | Prompt core |
|---|---|---|
| 1 | 5s | Extreme wide, tiny figure crossing a grey crater plain, low sun, static |
| 2 | 3s | Medium close-up, helmet reflection of dead panel lights, 85mm |
| 3 | 2s | Insert, gloved hand on a corroded connector, macro |
| 4 | 4s | Medium, static, amber light stutters across helmet visor |
| 5 | 2s | Close-up, sparks die, slow push in |
| 6 | 6s | Medium wide, slow pull back, she sits beside the beacon |
| 7 | 4s | Wide, distant point of light on the horizon, static |
Every one of those prompts carries the same palette line ("desaturated grey-blue with a single amber practical") and the same wardrobe anchor ("white hard-shell suit with a scorched right shoulder panel"). That repetition is what makes seven separate generations read as one film.
A Repeatable Production Workflow
Once you have a method, the workflow compresses. Here is a sequence that scales from a solo creator to a small team.
Stage one: pre-production
Write the logline, beats, and shot list. Build the character and world bible. Generate hero reference images for each character and location. Decide which model handles which shot function. Estimate durations and total runtime. Nothing here involves motion generation, and it is where most of the quality is decided.
Stage two: generation
Generate in shot order, not in order of excitement. Approve one shot before moving to the next so drift does not compound. Keep a simple log with model, seed, prompt version, and approval status for each shot. When a sequence breaks, that log tells you exactly where it happened.
Stage three: assembly
Edit to the durations you planned, then adjust. Normalise the grade across all clips, add grain or texture if your models differ stylistically, and only then cut. Sound design — ambience, a single sustained tone, minimal music — does more for perceived production value than another generation pass ever will. Try a version with only ambience and no music at all; it is often the stronger cut.
Stage four: review against the spine
Read your logline again and watch the finished piece. If an outside viewer cannot restate the logline after one viewing, a shot is doing aesthetic work instead of narrative work. Cut it. Shorter and clearer almost always wins.
Common Prompting Mistakes and How to Fix Them
Overloaded prompts. Four subjects, three actions, and a camera move in one shot. Fix: one action, one subject group, at most two camera ideas.
Paraphrasing continuity. "Green jacket" in shot two, "olive coat" in shot five. Fix: copy and paste the bible phrase every time.
Style before story. Choosing a cinematic look before knowing what happens. Fix: lock the beats first; style is a filter applied to a finished structure.
Generating in random order. Twenty unapproved shots and no idea which are usable. Fix: sequential approval with a generation log.
Ignoring runtime. Twelve shots at eight seconds each is a ninety-six-second film, not the forty seconds you imagined. Fix: sum durations before you generate.
Fixing drift with more generations. Re-rolling does not solve continuity. Fix: compare against the reference and change one variable.
No negative constraints. Unwanted watermarks, extra limbs, or modern objects creeping into a period piece. Fix: a short, consistent avoid list in every prompt.
FAQ
Do I need a different prompt style for each model?
The five building blocks stay the same. What changes is emphasis: some models respond strongly to lens and lighting language, others need simpler, more literal sentences. Test your template once per model and adjust phrasing, not structure.
How long should a single prompt be?
Long enough to be unambiguous, short enough to read in fifteen seconds. In practice, 30 to 60 words of dense description usually outperforms both a five-word tag cloud and a 200-word paragraph.
Should the prompt be written in first person or third person?
Third person, present tense, descriptive. "She reaches for the switch" rather than "I want a shot where she reaches." Instructions about your intent belong in your notes, not in the prompt.
What if my character changes face between shots?
Generate a hero portrait first, use it as an image reference for every subsequent shot, and repeat the same facial description word for word. If drift persists, keep that character in a tighter shot range — wide shots are less revealing and more forgiving.
Is storyboarding necessary if I have a shot list?
Not strictly, but rough thumbnails or even stick figures make duration decisions obvious in a way that text does not. Five minutes of sketching saves an hour of generation.
How do I handle dialogue?
Treat dialogue as a separate layer. Generate the visual with a performance that reads correctly silent — reactions, posture, eyelines — then add voice in post. Lip-sync generation is improving but still unreliable for anything beyond short lines, and a well-cut reaction shot with voice-over often feels more professional.
What is the fastest way to improve at this?
Recreate a thirty-second scene you already love. Copy its shot sizes, durations, and lighting choices into your template. Reverse-engineering a sequence you admire teaches more about pacing and camera language in one afternoon than a month of free experimentation.
The through-line in all of it is simple: the script is the product, and the prompt is the manufacturing specification. Models will keep changing, cameras will keep getting better, and the tool you used last year will look dated. What does not expire is a clear logline, a shot list with durations, a locked bible of visual anchors, and the discipline to write camera language instead of adjectives. Master that and every new generation model becomes an upgrade rather than a restart.


