Start With Intent, Not Aesthetics
Most first attempts at a video prompt describe a beautiful still image. "A woman in a red coat walking through a rainy neon street, cinematic lighting, ultra detailed." The clip that comes back is often gorgeous and completely useless, because nothing happens, nothing connects to the next shot, and the viewer has no reason to keep watching. Beauty is not a story.
A storytelling prompt answers three questions before it ever mentions a lens: who wants something, what is in the way, and what changes by the end of the shot? When you can answer those in one sentence, the visual details stop being decoration and start being evidence. The rain is no longer "cinematic" — it is the reason she cannot hear the footsteps behind her.
This guide is a practical workflow for writing prompts that behave like direction rather than description. It covers prompt structure, shot-list planning, consistency across multiple generations, model-specific translation, and fixes for the failures that show up most often. Everything here works whether you generate in a browser tool, through an API pipeline, or inside a desktop editor with a generative plugin.
The Five Layers of a Cinematic Prompt
A reliable prompt has five layers, ordered from meaning to mechanics. Beginners usually write layer three first and never write layer one. That inversion is the single biggest cause of unsatisfying results.
Layer 1: Intent
State the dramatic function of the shot in plain language before any visual description.
Weak: "Cyberpunk alley, dramatic lighting."
Stronger: "A courier realizes she has been followed and decides to stop running."
The second version tells the model — and, just as importantly, tells you — what the shot has to accomplish. Intent also gives you a test for the output: if the clip does not show a decision, it failed, no matter how good the fog looks.
Layer 2: Subject specifics
Name the subject precisely: age range, build, wardrobe with materials and colors, emotional state, and what they are holding. Vague subjects force models to guess, and guesses drift between shots. "A tired nurse in scrubs" becomes a different person in every generation. "A broad-shouldered nurse in her forties, navy scrubs, hair tied back, clutching a paper cup she never drinks from" gives you something repeatable.
Attach one or two physical anchors — a scar, a bright scarf, an unusual jacket — and reuse that exact wording in every prompt for that scene. Anchors are cheaper than consistency tooling and often work better.
Layer 3: Environment, light, and time
Now describe the world. Three variables matter most: time of day, weather, and the direction of the dominant light source. Light direction is the most neglected and the most powerful. "Backlit by a low sun" and "lit from above by a flickering fluorescent" produce entirely different emotional registers from the same set.
Keep the palette restrained. Two dominant colors plus one accent is enough for most scenes, and it prevents the muddled, over-saturated look that comes from stacking ten color adjectives.
Layer 4: Camera and lens
Camera language is where prompt writers earn their keep. Specify four things: shot size, angle, movement, and lens behavior.
- Shot size: extreme wide, wide, medium, close-up, extreme close-up.
- Angle: eye level, low, high, over-the-shoulder, overhead, Dutch tilt.
- Movement: static, slow push in, pull out, pan, tilt, handheld follow, crane up, orbit.
- Lens: wide-angle with mild distortion, normal 50mm feel, long lens with compressed background, shallow depth of field, deep focus.
One movement per shot. Two movements confuse the model and produce mushy motion.
Layer 5: Motion and sound cues
Describe the primary action with a clear beginning and end, then note secondary motion like hair, fabric, smoke, or rain. When your pipeline supports audio, add a short sound description: an ambient bed, one diegetic effect, and a musical tone. Even when audio is not generated, writing it clarifies pacing.
A full example combining all five layers:
"Intent: a father hiding bad news. Subject: a heavyset man in his fifties, gray stubble, olive work jacket, forcing a smile. Environment: cramped kitchen at night, single warm bulb above the table, rain on the window. Camera: medium shot, eye level, slow push in over six seconds, normal lens, shallow focus on his eyes. Motion: he sets down a fork, exhales, looks up. Sound: rain, refrigerator hum, one low piano note."
That prompt is boring to read and produces something that works.
Build the Shot List Before You Write Prompts
Generating one clip at a time invites drift. Before you open any tool, write a shot list on paper or in a plain text file. Eight to fifteen entries is plenty for a two-minute piece.
Each entry needs five fields: shot number, dramatic function, shot size, camera movement, and the one thing that changes. Keep the function column honest. If a shot's function is "establish mood," ask whether the previous shot already did that. Redundant shots are the most common reason AI films feel long at ninety seconds.
Group the list into sequences of two to four shots that share location and light. Sequences give you natural boundaries for testing: if a sequence works, you can move on; if it does not, you have a small, fixable problem instead of a whole film to redo.
Finally, note the connection points between shots — the match on action, the eyeline, the object that carries over. These connections are where continuity either holds together or falls apart, and they are much easier to plan than to repair.
Write Prompts That Survive Model Translation
Different engines interpret the same sentence differently. Some are tuned for photoreal texture and reward concrete nouns; others favor stylization and respond to art-historical references. You do not need a separate prompt for every model, but you do need a stable core with swappable parts.
Use a three-part template:
- Locked core — subject with anchors, wardrobe, environment, light direction. This never changes within a sequence.
- Variable block — shot size, angle, movement, and the action beat. This changes per shot.
- Style tag — a short, consistent phrase describing the overall look.
When you move from one engine to another, keep the locked core verbatim and rewrite only the style tag. Engines that respond well to technical language get lens and lighting terms. Engines that respond to mood get emotional and material adjectives instead: "damp wool," "late afternoon haze," "quiet dread." Try both once and keep notes on which phrasing your chosen tool obeys.
Two practical rules save hours. First, front-load the subject: the earliest tokens get the most attention. Second, remove hedge words like "slightly," "somewhat," and "maybe." Models treat them as noise, and they dilute the terms around them.
Keep Characters and Props Consistent Across Shots
Consistency is the hardest part of multi-shot AI video, and it is mostly a writing problem before it is a tooling problem.
Write a character sheet with a fixed descriptor string for each person: age, build, hair, one wardrobe item with color and material, one distinguishing feature. Copy that string, do not paraphrase it. "Dark green canvas jacket" and "forest-colored coat" will produce two different garments.
For props that matter to the plot, do the same. A brass key, a cracked phone screen, a red thermos — each gets a fixed phrase. Reuse the phrase the same way you reuse an actor.
When your tool supports reference images, stills, or character slots, use them in addition to the written anchors, not instead of them. The image sets the visual baseline; the text keeps it from wandering. When your tool does not support them, lean harder on anchors and choose wardrobe with strong silhouettes — a wide-brimmed hat or an asymmetric coat reads as the same person even when facial details shift.
Accept small variations. Audiences forgive a slightly different nose between shots; they do not forgive a character who changes age, hair color, and outfit simultaneously.
A Camera Language Cheat Sheet
Keep this list next to your shot sheet and pick deliberately rather than decoratively.
| Beat you want | Shot size | Movement |
|---|---|---|
| Establish scale | Extreme wide | Slow crane or drift |
| Introduce a place | Wide | Static or gentle pan |
| Show a decision | Medium | Slow push in |
| Reveal inner state | Close-up | Static, shallow focus |
| Signal danger | Medium | Handheld follow |
| End a scene | Wide or close-up | Pull out |
Movement should have a motive. Push in when a character commits to something. Pull out when they are left alone with the consequences. Hold still when the audience needs to read a face. If you cannot justify a camera move in one clause, delete it and generate a static shot instead — it will look better and cost you less time.
The Directorial Pass: Editing What the Model Gives You
Generation is not the last step; it is the first draft. Treat each clip as footage from a shoot you did not fully control.
Watch each clip three times with a different question in mind. First pass: does the action read? Second pass: does the light match the neighboring shot? Third pass: where is the frame most alive?
Then do the trimming pass. Cut the first and last few frames of most clips, since motion often ramps in and out awkwardly. Trim on movement so cuts land on action rather than after it. Add a short dissolve only when two shots share a location but not a beat.
If a clip fails, do not immediately regenerate blindly. Change one variable: reduce camera movement, shorten the requested action, or simplify the background. Single-variable iteration is how you learn what your tool actually does.
Troubleshooting the Failures You Will Actually Hit
The subject morphs mid-clip
Usually caused by too much simultaneous action or by contradictory descriptors. Cut the action to one beat, remove secondary characters from the prompt, and shorten the clip length. If it persists, add an anchor that is visible throughout — a bright jacket, a held object — so the model has a stable reference point.
Motion is mushy or smeared
This typically means you asked for two movements at once or described an action with no clear end state. Rewrite the action as a single motion with a destination: "she turns toward the door and stops" rather than "she moves emotionally."
The look drifts between shots
Style drift comes from rewriting the style tag per shot. Freeze the tag and copy it word for word. Also check your lighting descriptions; three different light directions across one scene will read as three different rooms.
Everything looks like stock footage
Stock-footage energy comes from generic subjects and unmotivated camera moves. Give the subject a specific, slightly inconvenient detail — a coffee stain, a bandaged hand, a mismatched shoe — and choose a camera angle a real crew would find awkward to rig.
The clip is technically fine but boring
That is a writing problem, not a generation problem. Return to layer one and ask what the character wants in this shot. If the answer is nothing, the shot does not need to exist.
A Ninety-Minute Practice Loop
Consistency improves fastest with deliberate repetition. Block ninety minutes and run the same loop three times with different material.
Minutes 0–15: write a four-shot sequence with a clear dramatic turn. Minutes 15–35: generate each shot twice with one variable changed. Minutes 35–55: assemble a rough cut and note where continuity breaks. Minutes 55–75: return to the prompt sheet and fix the weakest shot. Minutes 75–90: watch the cut with sound off, then describe what you understood.
Track results in a simple log: prompt, tool, settings, what worked, what failed. After a few weeks, the log becomes a personal reference that is far more valuable than any general prompt list, because it reflects the specific engine and style you actually use.
FAQ
How long should a single AI video prompt be? Long enough to cover the five layers and no longer. Most effective prompts land between 40 and 90 words. Beyond that, later clauses start competing with earlier ones for attention.
Should I write prompts in the language of the film's dialogue? Write in whichever language you are most precise in. Precision matters more than locale, though you should keep terminology consistent within a project so your style tag does not shift meaning.
Do I need a separate prompt for every model I try? No. Keep the locked core identical and adapt only the style tag and the technical vocabulary the engine understands.
How do I handle dialogue in AI video? Generate the visual performance and add dialogue in post. Time the line to the performance rather than trying to make the model speak, and keep shots shorter when a line carries the scene.
What is the fastest way to improve? Recreate a scene you already love from a real film, shot for shot, in your own words. You will learn more about framing, pacing, and continuity in one exercise than in a month of random generation.
Is storyboarding still worth it? More than ever. The storyboard is where you solve problems cheaply. Every minute spent planning saves several minutes of regeneration and avoids the slow, demoralizing drift that comes from improvising shot to shot.
The through-line across all of this is simple: direct first, describe second. Prompts are instructions to a crew that has never met you and cannot ask questions. The clearer your intent, the fewer questions it has to invent answers to — and the closer the footage lands to the film you saw in your head.


