Most disappointing AI video comes from a capable model and a weak plan. The prompt was vague, the camera had no reason to move, the jacket changed color between shots, and the final edit had to be rescued with music and speed ramps. None of those problems came from the generator. They came from skipping the part of filmmaking that has always mattered most: shot design.
This guide outlines a repeatable workflow for designing shots before you generate them, translating that design into prompts, holding characters and locations stable across clips, and cutting the results into something that reads as a scene rather than a demo reel. Tool names change every few months. The craft underneath does not.
Why Shot Design Matters More Than the Model You Choose
Every new video model raises the ceiling on realism, and every new model tempts creators to believe that better output is only a prompt away. It is not. A prompt like "cinematic shot of a woman walking through a rainy street" produces something watchable, but it does not produce a scene. Scenes are built from relationships between shots: a wide establishes where we are, a medium two-shot shows who is in the room, a close-up tells us what matters, an insert reminds us what is at stake. Remove that architecture and you get a sequence of attractive fragments with no argument and no emotion.
Shot design also solves a very practical generation problem: specificity. Diffusion and transformer video models respond to constraints. When you decide in advance that a shot is a 35mm over-the-shoulder with the camera slowly pushing in from hip height, you naturally write a prompt detailed enough to produce controlled output. Design is not decoration. It is prompt engineering with a reason behind it.
Finally, shot design is what makes AI footage editable. If every clip is a beautiful wide shot, an editor has nothing to cut between and no way to build rhythm. Coverage — the deliberate variety of angles, sizes, and moments — gives the edit choices. You are not generating clips; you are generating options.
The Cinematic Vocabulary That Translates Into Prompts
You cannot direct what you cannot name. Before prompting, get comfortable with the vocabulary that models actually respond to, because most of it maps directly onto prompt language.
Shot sizes and what each one does
- Extreme wide / establishing: Shows geography and scale. Use it to open a scene or reset location. Prompt as "extreme wide shot, subject tiny in frame, vast landscape."
- Wide / full shot: Shows the body and its relationship to the space. "Full body shot, subject centered, wide framing."
- Medium shot: The workhorse. Waist-up framing carries dialogue and gesture. "Medium shot, waist-up, subject on the left third."
- Medium close-up: Chest-up. Intimacy without losing context. "Medium close-up, chest-up framing."
- Close-up: Face fills the frame. Emotion and decision. "Close-up, face fills the frame, eyes in focus."
- Extreme close-up / insert: Detail on an object or body part. "Macro insert of hands tightening a knot."
- Over-the-shoulder: Puts two characters in a relationship and implies a point of view. "Over-the-shoulder shot from behind the older man, facing the young woman."
Models frequently ignore abstract labels like "medium shot," so pair each term with a physical description of framing. "Waist-up" and "face fills the frame" are harder to misinterpret than a film-school term.
Lens language, depth, and perspective
Lens choice changes meaning as much as framing does. An 18–24mm lens exaggerates space and makes rooms feel larger and more threatening. A 35mm lens feels documentary and natural. A 50mm lens approximates human perspective and works well for neutral coverage. An 85mm lens compresses the background and flatters faces, which is why portraits use it. Macro lenses reveal texture. Anamorphic optics add wide flares and oval bokeh that read as "cinema" to most viewers.
Write lens and depth into the prompt: "shot on 85mm, shallow depth of field, background melts into soft bokeh," or "24mm wide lens, deep focus, foreground and background both sharp." Depth of field is one of the strongest tools for directing attention, and it is one of the easiest things to request.
Camera movement as emotional punctuation
Static frames feel observational and calm. A slow push-in builds pressure and pulls the audience toward a thought. A dolly out isolates a character. A handheld shot signals urgency, immediacy, or instability. A gimbal glide feels smooth and slightly inhuman, which suits dream sequences and product reveals. A crane or drone move delivers revelation. An orbit adds ritual and drama around a subject. A whip pan implies speed and cuts well into another whip pan.
One rule saves a lot of pain: one primary movement per clip. Models blend conflicting instructions into mush, so choose push-in or orbit, never both, and describe the movement's speed as well as its direction — "slow push-in," "gentle drift left," "fast handheld follow."
Build a Shot List Before You Touch a Prompt Box
A shot list is the cheapest part of production and the highest-leverage. Write it in a plain spreadsheet or document with six columns: shot number, description, purpose, size and movement, target duration, and audio notes.
A three-column starting point
If a full list feels heavy, start with three columns: what we see, what it does for the story, and how it should feel. A row might read: "Wide of the empty kitchen at dawn / establishes that she is alone now / quiet, cold, still." That single row already contains the framing, the emotion, and the pacing instruction. Turn it into a prompt later, not now.
Coverage strategy for AI shots
For each scene, plan a minimum of three shots: one establishing shot to orient the viewer, one performance shot that carries the emotion, and one detail shot that adds texture or information. Add cutaways — hands, feet, weather, objects, backgrounds — because they are cheap insurance in the edit when a clip does not hold up. Because generation is non-deterministic, plan to produce two to four variants per shot and treat the shot list as a target rather than a contract.
Group your list by location and wardrobe state before you generate anything. This single step prevents the most common continuity disaster in AI filmmaking: generating scenes in random order and discovering that the character's hair is wet in one shot and dry in the next.
Turn the Shot List Into Director-Grade Prompts
A shot list that never becomes a prompt is just a wish. The translation step is where most creators lose control, usually because they write one long sentence with everything crammed together.
The five-part prompt structure
- Subject and wardrobe: Who is in frame, wearing what, in what condition. "A fisherman in a faded yellow raincoat, hood down, salt-stained sleeves."
- Action and beat: One action, one moment. "He lifts a rope coil and steps onto the pier."
- Camera: Size, angle, lens, movement. "Wide low-angle shot, 28mm, slow dolly right."
- Light and atmosphere: Time of day, source, weather, air. "Overcast dawn light, sea mist, wet reflective stone."
- Style and grade: Texture and color. "Documentary realism, muted teal and grey grade, natural grain, 2.39:1 framing."
Put the subject and action first, camera second, atmosphere and style last. Front-loading what matters keeps the important details intact even when a model truncates or reweights your text.
Writing negative instructions that actually help
Long negative lists rarely work in prose. Keep them short and concrete: "no text overlays, no duplicated limbs, no camera shake, no slow-motion." Many interfaces handle negatives as separate parameters, so check whether your tool exposes a dedicated field before padding your prompt with "no" clauses. Also avoid style buzzwords that carry no visual instruction — "epic," "stunning," and "masterpiece" consume attention without changing a single pixel.
Example: one scene at three shot sizes
Same scene, same character, three deliberate framings:
- Wide: "Extreme wide shot, a lone fisherman in a faded yellow raincoat walks along a wet stone pier at dawn, 24mm lens, static camera at low angle, overcast light, sea mist, muted teal and grey grade, natural grain."
- Medium: "Medium waist-up shot, the same fisherman in a yellow raincoat pulls a rope coil across the pier, 50mm lens, slow tracking shot moving with him, overcast dawn light, wet stone, muted teal and grey grade."
- Close: "Close-up, the fisherman's weathered face fills the frame as he looks toward the horizon, 85mm lens, shallow depth of field, soft mist bokeh behind him, overcast dawn light, muted teal grade."
Notice what stays constant and what changes. Wardrobe, location, time of day, and grade are identical across all three; only framing, lens, movement, and the specific action change. That discipline is what makes three separate generations feel like one scene.
Lock Character and Location Consistency
Consistency is not a prompt trick. It is a production system built from references, descriptors, and order of operations.
Reference conditioning and character sheets
Build a character sheet before generating any scene: three to five images of the same face from different angles, plus a wardrobe shot. Then use image-to-video or reference-conditioned modes rather than relying on text alone. Text descriptions of a face drift; a reference image keeps the model anchored to actual pixels.
Keep a short, fixed descriptor block for each character and reuse it verbatim in every prompt — same words in the same order. Changing "silver-streaked beard" to "grey beard" between shots can shift age, jawline, and lighting. Fixed language plus fixed references is the closest thing to a reliable consistency pipeline.
Continuity details people forget
- Wardrobe layers: Jackets on or off, sleeves rolled up, gloves present.
- Hair and skin state: Wet, dry, tied back, bloodied, sweating.
- Props in hand: Rope, cup, phone, weapon — and which hand holds it.
- Screen direction: If a character walks left-to-right in the establishing shot, keep them moving left-to-right in the next one.
- Time of day and weather: Cloud cover and color temperature need to match across clips or the cut reads as a mistake.
- Extras and background life: Population density in a street shouldn't jump from empty to crowded.
Write these into your shot list as a checklist column. It takes two minutes and saves entire re-generation sessions.
Direct Camera Movement and Temporal Pacing
Movement is meaning. Map it deliberately instead of asking for "dynamic camera" on everything.
Match movement to emotion
- Realization: slow push-in on a still frame.
- Isolation or defeat: dolly out, subject shrinking in frame.
- Urgency: handheld follow, slight instability.
- Reveal or ritual: crane up, orbit, or a slow arc.
- Unease: a slow pan that leaves the subject slightly off-center.
- Speed and transition: whip pan, which cuts well into another whip pan.
Runtime, beats, and cut rhythm
Generate clips in short units — roughly four to eight seconds — and give each one a single completed beat. Asking a model for two actions in one clip usually produces one action done poorly. Plan your cut rhythm in advance: dialogue scenes often cut every three to five seconds, action sequences every one to two, and emotional beats can hold much longer than instinct suggests. If a line of dialogue takes four seconds to say, the shot covering it should be about that long, plus a beat of reaction.
Direct With Sound in Mind
Silent generation is a trap. Plan audio while you plan shots, not after.
Start with three layers: dialogue and voices, ambience, and foley or score. Most teams generate picture first, then add voices with a text-to-speech or voice-cloning tool, then layer an ambience bed and specific sounds — footsteps on wet stone, rope creaking, gulls. That third layer is what convinces an audience the image is real.
When you write dialogue, write short lines with clear rhythm. Long monologues are difficult to match to lip movement and easy to over-perform. If your model supports native audio, test whether it respects speech timing; if not, generate picture without audio and sync later using a dedicated lip-sync pass. Aspect ratio matters here too — deliver in the ratio you plan to edit in, because cropping after generation loses composition you carefully designed.
Edit for Rhythm, Then Fix Continuity
Assembly and the rough cut
Lay shots in story order, then cut for rhythm rather than for completeness. Cut on motion whenever possible — a door closing, a head turning, a hand moving — because movement hides the edit. Use L-cuts and J-cuts with your ambience track to smooth transitions, letting sound from the next scene arrive before the picture does. If a shot is beautiful but slows the scene, cut it. The shot list exists to serve the sequence, not the other way around.
The pre-render checklist
Before you generate the next batch, confirm: framing language is explicit, one movement per clip, character descriptor block pasted verbatim, wardrobe and time of day match the previous shot, screen direction preserved, aspect ratio set, target duration set, and audio plan noted. Nine checks, ninety seconds, far fewer wasted generations.
Common Mistakes and How to Fix Them
| Mistake | Why it hurts | Fix |
|---|---|---|
| Every shot is a wide | No rhythm, nothing to cut between | Force a three-shot minimum per scene |
| Stacked camera moves | Motion blurs into chaos | One primary movement per clip |
| Vague subject description | Character drifts between clips | Fixed descriptor block plus reference images |
| Prompt soup | Model ignores half the instruction | Use the five-part structure |
| No cutaways | No recovery when a shot fails | Generate hands, objects, weather inserts |
| Buzzword style prompts | Changes nothing visually | Replace with lens, light, and grade specifics |
| Generating in random order | Continuity breaks everywhere | Generate grouped by location and wardrobe state |
| Asking for a full scene in one clip | Weak action, no coverage | Generate short beats, assemble in the edit |
FAQ
Do I need a storyboard, or is a shot list enough?
A shot list is enough for most creators. Storyboards help when you need to communicate framing to collaborators or when a scene depends on precise spatial relationships. If you can describe the frame in one sentence, you can skip drawing it.
How long should each AI-generated clip be?
Four to eight seconds covers most beats comfortably. Shorter clips are easier to control and cheaper to iterate on; longer clips tend to introduce drift in faces, hands, and background detail. Build long scenes from short pieces.
Why does my character's face change between shots?
Usually because the model is working from text alone. Use reference images, keep the descriptor block identical across prompts, and generate all shots of one character in a single session so environmental and style settings stay consistent.
How many generations should I plan per shot?
Two to four variants is a realistic baseline for a usable take. Complex action or crowded frames may need more. Budget your time by shot complexity, not by total shot count.
What aspect ratio should I use?
Match your delivery platform and edit in that ratio. Widescreen suits cinematic storytelling and landscapes; vertical suits short-form feeds and faces. Decide before you generate, because reframing after the fact destroys compositions you designed on purpose.
Can I change camera movement after the video is generated?
Rarely in a satisfying way. Some tools offer limited motion controls or camera paths on generated footage, but the results are soft compared to designing the movement into the original prompt. Treat movement as a generation-time decision.
Do I need professional editing software?
No, but a real timeline helps. Any editor that supports multiple tracks, precise trimming, and audio mixing will do. The edit is where shot design pays off — without the ability to trim to the frame, even a well-designed sequence will feel loose.


