Start With the Story, Not the Settings
Every generation tool now ships with a wall of sliders: model choice, aspect ratio, motion strength, seed, style reference. It is tempting to open the interface and start experimenting. That instinct produces beautiful fragments and unusable films. The projects that survive contact with an audience are the ones where the narrative problem was solved first, in plain language, on a page.
Think of the division of labor the way a real production does. A director does not operate every department; they decide what the scene needs and communicate it precisely. In AI video, your job is the same. The model handles rendering, but you own intent: who wants what, why now, and what the audience should feel when the cut lands.
A useful test before you generate anything: can you describe each shot in one sentence that includes a subject, an action, and a change? "Mira hesitates at the door, then steps through" is a shot. "Cinematic woman, dramatic lighting, 4K" is wallpaper. The first has a before and after state; the second has no reason to exist.
Constraints are the second half of that test. Pick a single visual premise — one palette, one lens family, one lighting logic — and defend it across every scene. Consistency reads as competence; variety without discipline reads as noise. When you keep the premise stable, small deviations, such as a warmer key light in the reconciliation scene, become meaningful instead of random.
Finally, decide the delivery format before you write the shot list. A vertical short rewards tight framing and fast escalation. A horizontal explainer rewards establishing shots and readable text. The same script deserves different coverage in each case, and reconciling that after generation is expensive.
Pre-Production That Survives Generation
Logline, premise, and the promise to the viewer
Write a logline of 25 to 40 words that names protagonist, desire, obstacle, and stake. Then write a second sentence describing the emotional promise: "A quiet revenge story where the tension comes from politeness, not violence." That second sentence is what you hand to an AI assistant or use as the benchmark when reviewing outputs. If a generated shot contradicts the promise, it goes.
Beat sheets and scene cards
A beat sheet is a list of irreversible changes, not a list of events. For a 90-second piece, six to nine beats is usually right. Each beat gets a scene card with four fields: location, time of day, who is present, and what changes by the end. Cards are cheap to reorder; generated footage is not. Rearranging cards on a table reveals pacing problems that no amount of rendering can fix.
The shot list as a contract
For each scene card, write two to six shots. Note framing (wide, medium, close), subject position, camera behavior (static, push, pan, handheld), and the transition out. This list becomes your generation queue and your review checklist. When a shot is hard to produce, the list tells you whether you actually need it, or whether you were covering a beat you could deliver with a reaction instead.
How an AI Director Reads Your Script
AI directing tools — sometimes framed as agent directors or story copilots — are most useful as structural editors rather than prompt polishers. They analyze input text for narrative shape and suggest coverage. Understanding what they look at helps you feed them better material.
Conflict mapping
The assistant looks for who opposes whom and where the pressure peaks. If your script has no explicit opposition, it will either invent one or flatten the tension. Fix it at the source: give the antagonist a concrete tactic, not a personality trait.
Emotional tone curves
Tools assemble an emotional profile per scene — calm, wary, urgent, relieved — and check whether the transitions are motivated. This is where AI feedback is genuinely valuable: it flags a jump from grief to slapstick that your outline forgot to bridge.
Structural gap detection
Expect suggestions like "the midpoint has no reversal" or "the ending resolves a question the audience never asked." Treat these as diagnostic prompts. You can accept a suggestion, replace it with your own fix, or consciously override it — but you should be able to explain the override in one sentence.
The practical use of these tools is negotiating coverage, not writing your story. Ask for three alternative shot approaches per beat, then choose the one that keeps your visual premise and cuts cleanly to the next beat.
Shot Design Fundamentals That Transfer to AI
Framing and psychological angle
Framing is an argument about power. Low angles make a subject dominant; high angles make them exposed. Eye-level, slightly off-center framing reads as observational and lets the audience decide. In AI generation, framing must be stated explicitly because models default to a flattering, centered composition. If you want a subject small in a large room, say so, and specify what occupies the empty space.
Lens and depth cues
Lens language changes perceived intimacy. Wide lenses exaggerate space and movement; long lenses compress and isolate. Most video models respond well to "shallow depth of field, 85mm portrait compression" or "wide 24mm, deep focus, foreground clutter." Depth cues — foreground objects, atmospheric haze, layered silhouettes — do more for realism than any resolution bump.
Movement and cut rhythm
Camera movement should have a motivation: following a decision, revealing information, or withholding it. A slow push on a face during a hesitation is a different story beat than a handheld follow during a chase. When planning cuts, alternate shot sizes rather than angles of the same size; the mismatch creates momentum. For a 20-second scene, three shots with one deliberate movement beat usually outperform eight drifting clips.
Character and World Consistency Without a Studio
Continuity is the number-one complaint in AI video. Faces drift, jackets change color, rooms rearrange. Treat consistency as data, not luck.
Build a character sheet for every recurring subject: age range, build, hair, distinguishing marks, wardrobe with exact colors, and two or three phrases that capture posture and expression. Save two to four approved reference stills per character, taken from different angles, and reuse them in every generation. For locations, do the same: an empty plate of the room, the lighting direction, and a note on which surfaces are reflective.
Then define continuity rules that are checkable at a glance: which side of frame the subject enters from, what time of day each scene occurs, what props must be visibly present. Reviewers catch violations quickly when the rules are written down. Without written rules, every reviewer invents their own standard.
Handle wardrobe and props like a low-budget production: limit changes. One costume per character per act, one key prop per scene. Fewer variables means fewer failures, and the audience tracks story instead of continuity errors.
Keyframes, Style Frames, and Image Fusion
Keyframes are the bridge between storyboards and motion. Generate a still for the first and last frame of each shot, approve them, then animate between them. This is dramatically more controllable than describing motion in text, because static composition is where image models are strongest.
Style frames set the grammar for the whole piece: one frame that establishes palette, contrast, grain, and lens character. Approve it before producing anything else, then attach it as a visual reference for every subsequent generation. When drift appears mid-project, compare against the style frame — not against the previous shot, which may itself have drifted.
Multi-image fusion techniques let you combine a character reference, a location plate, and a style frame into a single coherent request. Keep the number of references small and mutually compatible. Three references that agree on light direction beat six references that each contribute a different mood. Label your reference images clearly in your project folder, and note which shot each one belongs to; the moment a project passes 30 assets, unlabeled files become the real bottleneck.
A Camera-Aware Prompt Template
Adopt a fixed prompt order so your team can read any generation request quickly:
- Subject and action in the present tense.
- Framing and lens: shot size, angle, focal length feel.
- Camera behavior: static, push in, pan left, handheld follow.
- Lighting and time: direction, quality, color temperature.
- Environment and depth: foreground, background, atmosphere.
- Style reference: palette, grain, film stock, era.
- Continuity anchors: wardrobe, props, character reference ID.
Example: "Mira, mid-30s, dark bob, charcoal coat, pauses at a glass door, hand on the handle. Medium close-up, slight low angle, 50mm. Slow push in, camera static otherwise. Late afternoon window light from camera left, soft with visible falloff. Foreground: blurred railing. Background: empty corridor, faint reflections. Muted teal-and-amber palette, fine 35mm grain. Continuity: charcoal coat, silver ring, reference CHAR-MIRA-02."
Keep negative constraints short and specific — flicker, extra fingers, warped text — rather than long lists of generic bans. And save prompts that work as named presets. A prompt library is the difference between a hobby and a repeatable pipeline.
Quality Control, Iteration, and Decision Criteria
The continuity pass
Review at normal speed, then again paused on cuts. Check identity, wardrobe, props, eyeline, and light direction. Rate each issue as fixable in post, fixable with a regenerate, or fatal to the shot. Only regenerate fatals first; otherwise you will burn days polishing shots that may be cut.
The motion pass
Watch for unnatural limb motion, morphing backgrounds, and speed changes. If a shot is 80 percent good but has one broken second, consider trimming rather than regenerating — editing is cheaper than generating. For dialogue-adjacent shots, check that mouth and head movement do not fight the audio.
Converging instead of perfecting
Set an acceptance bar before you review. For a social short, "passes on a phone screen at normal speed" is often enough. Track how many attempts each shot takes. If a shot exceeds five attempts, the problem is usually the concept, not the settings — rewrite the shot as two simpler shots.
Always keep a "maybe" bin. Shots that fail the primary intent often work perfectly as inserts, transitions, or background plates.
Common Mistakes and How to Avoid Them
- Prompt-only storytelling. Writing cinematic adjectives without a character goal. Fix: one sentence of intent per shot.
- Changing style mid-project. Each new reference pulls the palette. Fix: lock a style frame and attach it to every request.
- Too many shots per beat. Over-coverage makes editing a puzzle. Fix: two to four shots per scene, alternate sizes.
- Ignoring the vertical frame. Center-weighted compositions lose their subject under interface overlays. Fix: keep faces in the middle band and leave the top and bottom free.
- Regenerating instead of re-editing. Fix: exhaust trims, speed ramps, and sound design before new generations.
- No naming convention. Fix: name every asset with project, scene, shot, and version, plus a shot list that links to files.
- Skipping sound. Fix: add temp music and ambience before final review; rhythm changes what "works" means.
Frequently Asked Questions
Do I still need a script if the AI can generate from a prompt?
Yes. A script or beat sheet defines intent and order. Even a 15-second piece benefits from written beats, because they let you evaluate shots against a purpose rather than a vibe.
How many reference images per character are enough?
Two to four well-lit, varied angles usually hold identity. More references help only if they agree on lighting and wardrobe.
Should I generate first or storyboard first?
Storyboard or keyframe first for anything with more than one shot. Text-to-video is excellent for exploring texture and mood, not for locking coverage.
How do I keep style consistent across 20 shots?
Lock a style frame, use a fixed prompt order, keep the palette small, and review in batches rather than shot by shot.
When should I switch models?
Switch when a specific capability is missing — text rendering, longer takes, stronger subject preservation — not because a leaderboard changed. Test the same shot across two tools and compare against the acceptance bar you set.
How long should a scene be?
For short-form, three to eight seconds per shot and 15 to 40 seconds per scene keeps attention. Longer scenes need escalating information, not slower pacing.
Is it worth using an AI assistant for structure?
Yes, as a second reader. It is good at spotting unmotivated transitions and thin conflict. It is bad at knowing your taste, so treat its notes as options.
What is the fastest way to improve at shot design?
Rebuild a scene you admire shot by shot in writing: size, angle, movement, and what changes. Do that for ten scenes and your instinct for coverage sharpens faster than any settings experiment.




