Start With the Shot, Not the Prompt
Most people open a text-to-video tool and immediately start typing adjectives. That is the wrong first move. Cinematic results come from deciding what a shot has to do before you describe how it looks. A scene is a chain of small jobs: establish a place, introduce a face, reveal a detail, land an emotion, release tension. When every generated clip is assigned exactly one job, editing becomes trivial and the sequence feels directed rather than assembled.
Think in shot sizes before you think in words. Wide shots carry geography and isolation. Medium shots carry behavior. Close-ups carry feeling. If you cannot say in one sentence what a shot contributes, delete it from the list — generation time is better spent on shots that earn their place.
There is a second reason to plan first. Generated motion is convincing but not obedient. It handles atmosphere, light, and continuous movement beautifully, and it struggles with precise physical interaction between people and small objects. A shot list that ignores this reality produces a folder of near-misses. A shot list built around it produces footage that cuts together on the first pass.
So the working order is: job, shot size, subject action, lighting, movement, format. The prompt is the last step, not the first.
How a Text-to-Video Model Reads Your Description
What the model actually weighs
A video generator works by starting with noise and progressively resolving it into moving images, guided by your text and by how every part of the frame relates to every other part across time. Practically, that means relationships matter more than keywords. “Chef, kitchen, dough, marble, morning, flour” is a pile of nouns. “A chef pressing dough on a marble counter in a sunlit kitchen” is a description the model can resolve consistently, because it knows which words modify which.
Grouping is the whole game. Cluster related information: subject plus wardrobe; lighting plus time of day; lens plus framing plus movement. When unrelated details sit next to each other, the model averages them, and you get the glossy, indistinct look that reads as generated to any viewer.
Strengths worth designing around
- Motivated, natural-looking light: window light, firelight, neon signage, overcast diffusion, headlights.
- Continuous camera motion: pushes, lateral tracking, cranes, gentle handheld drift.
- Atmosphere: rain, dust, smoke, snow, heat shimmer, fabric and hair movement.
- Single-subject action inside one coherent space.
- Strong stylistic range, from documentary realism to painterly surrealism and archival film emulation.
Weaknesses you should hide in the edit
- Precise choreography between two or more people.
- Legible on-screen text, signage, and logos.
- Very long takes where the action changes significantly mid-shot.
- Exact facial continuity across separate generations.
- Fine hand interaction with small props.
If a shot requires complex physical interaction, split it into two shots and cut between them. Audiences read two shots as one continuous action; they read one broken shot as a mistake.
The Five-Layer Prompt Formula
Weak prompts fail because layers get shuffled together. Build in five stacking layers instead, and change exactly one layer per iteration.
Layer 1 — Subject and action
One subject, one action, one place. “A pastry chef pressing dough on a marble counter in a small bakery” outperforms “chef, bakery, busy kitchen, many people, cooking.”
Layer 2 — Framing and lens
Shot size, focal length, angle. “Medium close-up, 50mm, eye level,” or “wide establishing shot, 24mm, slightly low angle.” Shot size sets emotional distance, lens sets spatial distortion, angle sets power dynamics.
Layer 3 — Lighting
Direction, quality, source. “Warm morning window light from camera left, soft shadows, gentle falloff” or “single overhead practical, hard shadows, cool ambient fill.”
Layer 4 — Camera movement
Exactly one verb. “Slow push-in.” “Lateral tracking left.” “Crane down from ceiling height.” Stacking moves produces floaty, unmotivated motion that reads as artificial no matter how beautiful the frame is.
Layer 5 — Texture, palette, format
“2.39:1 anamorphic framing, fine 35mm grain, muted amber and teal palette, shallow depth of field.” This layer is what separates a graded-looking shot from a raw render.
A finished prompt
A 40-year-old pastry chef with short grey hair and a flour-dusted apron presses dough on a marble counter in a small bakery. Medium close-up, 50mm lens, eye level. Warm morning window light from camera left, soft shadows, gentle falloff. Slow push-in. 2.39:1 framing, fine film grain, muted amber palette.
Two to four sentences covers all five layers for most shots. If a prompt runs longer, you are usually repeating yourself or quietly introducing a contradiction.
Camera Language That Reads as Intentional Cinema
Certain moves are so embedded in film grammar that they read as deliberate even when generated. Others read as accidents. The difference is almost always specificity.
- Slow push-in. Tension and intimacy. Best for dialogue, realizations, and quiet beats. Phrase it plainly: “slow push-in, minimal drift.”
- Pull-back reveal. Widens context to close a scene. Strong for endings and punchlines.
- Lateral tracking. Follows a walking subject with documentary or thriller energy. Always name the direction.
- Crane up or down. Establishes scale. Keep it slow or it reads as a malfunctioning drone.
- Handheld follow. Energy and immediacy. Add “subtle sway, no shake” to keep it controlled.
- Short arc or orbit. Reveals a subject in three dimensions. Keep the arc at 20 to 30 degrees, because long orbits warp faces badly.
- Rack focus. Moves attention between foreground and background: “focus pulls from the foreground object to the subject's face.”
- Static wide with subject movement. The most underrated option. The camera holds, the subject crosses frame. Easy to generate, easy to cut, and it looks completely deliberate.
Matching movement to shot length
If a clip is under four seconds, movement gives it life. If it is longer, a locked-off or nearly static frame with a moving subject almost always looks more cinematic. Long generated moves tend to drift, and static frames hide that drift because there is no competing motion to compare against.
Name the direction, always
“Dynamic camera” is not a direction, it is a wish. “Lateral tracking right, subject centered” is a direction. Direction also protects continuity: if the subject exits frame right, the next shot should bring them in from frame left.
Lighting, Composition, and Format Decisions
Lighting language that changes output
Direction and hardness carry more weight than color names. Phrases worth keeping in rotation: “hard directional sunlight,” “overcast diffusion with no visible shadows,” “single practical lamp motivating the key,” “rim light separating subject from background,” “negative fill on camera right for contrast,” “warm bounce from a nearby wall.” Anchor everything with a time of day: “late golden hour,” “pre-dawn blue hour,” “midday with harsh overhead sun.”
Composition cues worth naming
Models default to centered, slightly flat framing. Counter it with depth: “foreground occlusion from a doorway, subject in the mid-ground, blurred street behind.” Other reliable cues include leading lines, frame-within-a-frame, subject placed on the left third, and deliberate negative space. Add “looking slightly off camera left” to avoid the dead-eyed stare; eyeline placement does more for believability than most people expect.
Aspect ratio, lens family, texture
Decide format before writing a single prompt, because it changes your framing instincts and your entire shot list.
- 2.39:1 — wide and anamorphic, more horizontal information. Ideal for landscapes, car interiors, and lonely streets.
- 16:9 — the safe delivery default and the format most clients expect.
- 9:16 — compose vertically from the start; keep faces higher in the frame with headroom for captions, and reduce camera movement, which reads as chaotic in a tall frame.
- 1:1 — useful for social cutdowns and product inserts that need a square crop.
Lens language tells a story on its own: 24mm for scale and environment, 50mm for neutral observation, 85mm for intimacy and compression. Texture language — fine grain, halation around highlights, slightly lifted blacks, gentle highlight roll-off — is what makes a clip feel graded rather than raw.
Continuity Across Shots: The Real Bottleneck
The character block
Write one paragraph describing your character, then paste it verbatim into every prompt that features them. Include age, hair, build, wardrobe, one distinguishing feature, and one recurring prop. Do not paraphrase between shots. Even a small wording change can drift a face. Where reference images are supported, generate a strong still first and reuse it as the visual anchor for everything that follows.
The look bible
Four to six sentences covering palette, contrast, grain, lens family, and era reference. Example: “muted teal and amber palette, low contrast in shadows, fine 35mm grain, soft highlights, 1970s European cinema feel.” Apply it to every prompt in the project. When two shots feel like they came from different films, the look bible was almost always dropped somewhere.
Screen direction and handles
Respect the 180-degree rule even though no physical camera exists. Track eyelines so two people generated separately appear to look at each other. Generate a second of extra material at the head and tail of every clip; editors need handles far more than they need tight clips. Two seconds of slack on a four-second shot costs almost nothing and can save an entire sequence in post.
From One-Page Script to Final Cut: A Working Process
Step 1 — Write a one-page script and a shot list
Six to ten shots per scene. For each, note shot size, subject action, and the emotional job it performs. Flag any shot that requires precise physical interaction and rewrite it as two shots now rather than discovering the problem later.
Step 2 — Build the character block and the look bible
Two short documents you will paste into every prompt. They take twenty minutes to write and save hours of regeneration.
Step 3 — Generate a hero frame before a hero shot
Iterate on stills until you find the image that defines the scene's look. Then describe that image back into motion. Stills converge faster, are quicker to judge, and give you a reference to measure every clip against.
Step 4 — Generate three to five variations per shot, changing one variable
First framing, then lighting, then movement. Keep a simple log: prompt version, variation number, what changed, what improved. Without a log you will repeat the same failure a week later and blame the tool instead of the wording.
Step 5 — Select and assemble
Choose takes that cut together rhythmically rather than takes that look best in isolation. Cut on motion: mid-movement cuts hide small imperfections, while cuts placed on static frames expose them.
Step 6 — Sound first, then grade
Build a scratch soundscape before judging the cut: room tone, footsteps, cloth movement, distant traffic, a low music bed. Silence is the single biggest reason generated footage feels fake. Then match contrast, black levels, and grain across shots. If one shot looks cleaner than the rest, add grain to it rather than stripping grain from everything else.
Step 7 — Mix generated and practical footage deliberately
Practical inserts — a real hand lifting a real cup, a real door closing — cut seamlessly with generated shots and take minutes to capture. Match lens language, grain, contrast, and movement style, and the two blend far better than most people expect on the first attempt.
A Worked Example: Six Shots at Night
This scene deliberately avoids crowds, dialogue, on-screen text, and complex hand interaction. It leans on atmosphere, movement, and cutting rhythm instead.
Shot 1 — Establish. “Rain-soaked city street at night, neon signage reflecting in puddles, empty crosswalk. Wide 24mm, slightly low angle. Practical neon as the only light source, deep shadows. Slow crane down. 2.39:1, fine grain, cool blue with warm neon accents.”
Shot 2 — Introduce. “A courier in a dark green rain jacket rides a bicycle through the street, water spraying from the rear wheel. Medium tracking shot, lateral tracking left, steady. Backlit by passing headlights, wet highlights. 2.39:1, subtle grain.”
Shot 3 — Detail insert. “Close-up of water droplets running along a bicycle handlebar, a gloved hand gripping it. 85mm, macro feel. Hard streetlight from above, high contrast. Static camera, minimal drift.”
Shot 4 — Reaction. “Over-the-shoulder shot of the courier pausing beneath a doorway light, breathing hard, rain dripping from the hood. Medium close-up, 50mm. Warm practical above, cold ambient behind. Static camera.”
Shot 5 — Emotional beat. “Close-up of the courier's face lit by a glowing phone screen, faint reflection in the eyes, rain on the glass behind. 85mm, shallow depth of field, no legible text on the screen.”
Shot 6 — Resolve. “Very wide shot of the courier riding away from camera into a lit intersection, growing small in frame. Static camera. City lights as background bokeh. 2.39:1, cool palette, heavy atmosphere.”
Why this sequence holds together: every shot contains one subject, one action, and one dominant light source. The two close-ups sit between wider shots, which masks continuity drift because the audience has no steady reference to compare faces against. Nothing depends on legible text or precise choreography. The whole scene lives on the exact qualities these models handle best.
Mistakes That Break Cinematic Results, and Their Fixes
- Overstuffed prompts. Fix: one subject, one action, one camera move. Delete every adjective that does not change the image.
- Contradictory lighting, such as a sunny overcast night. Fix: pick one time of day and one dominant source.
- Vague movement. “Dynamic camera” produces aimless drift. Fix: name the move and the direction.
- No format line. Fix: state aspect ratio and texture in every prompt, even when you are in a hurry.
- Dropping the character block between shots. Fix: paste it verbatim every time, including for inserts and background figures.
- Generating before deciding cut points. Fix: plan where cuts land so you know how much handle each clip needs.
- Judging clips in isolation. Fix: watch a rough assembly first, because a weak-looking clip often cuts beautifully in context.
- Skipping sound. Fix: always build a scratch soundscape before declaring that a shot failed.
- Over-correcting stabilization. Fix: rubbery motion looks worse than a slight handheld drift.
- Forcing long single takes for complex action. Fix: split into two shots and cut. Cheaper, faster, and more convincing.
Frequently Asked Questions
How long should a prompt be?
Long enough to cover the five layers, short enough to stay coherent. Two to four sentences is the sweet spot for most shots.
Can I keep one character consistent across many shots?
Yes, with discipline: identical character text in every prompt, a reference still as the anchor, consistent wardrobe and lighting, and more regeneration patience than a single shot would need. Close-ups drift first, so generate extra takes for those.
Should I write prompts in English if my project is in another language?
Video models are trained overwhelmingly on English descriptions, so English prompts tend to be more predictable. You can prompt in other languages, but translating the cinematography terms — lens, movement, lighting direction — into English usually improves accuracy. Keep dialogue and titles out of the visuals and add them in post.
How many variations should I generate per shot?
Three to five, with one variable changed at a time. Ten variations of the same flawed prompt teaches you nothing.
Can generated footage share a timeline with camera footage?
Absolutely, and that is where the technique becomes genuinely useful for client work. Match grain, contrast, lens family, and movement style, and the eye accepts the blend.
Do I need cinematography knowledge to get good output?
Not to begin, but it accelerates everything. Learning shot size, focal length, lighting direction, and movement vocabulary is the fastest route to better results — much faster than collecting prompt templates.
What about vertical video?
Compose for 9:16 from the start rather than cropping later. Keep subjects centered-high, favor medium and close shots, and reduce camera movement, which feels chaotic in a tall frame.
Is generation a replacement for practical shooting?
No. It is strongest for previsualization, atmosphere, establishing shots, and stylized inserts — anything that would otherwise need permits, weather, travel, or a large crew. Use it where speed and mood matter more than precise performance.
What is the fastest way to improve after a weak first attempt?
Rewrite one prompt using the five layers, change nothing else, and compare the two results side by side. Fixing prompt structure solves most problems before you touch any settings.



