Start With the Decision the Script Has to Force
Generating video has never been cheaper. Planning video has never been more valuable. When a single person can produce twenty shots in an afternoon, the bottleneck moves away from cameras, crews, and render farms and lands squarely on the question every script answers: what is this video supposed to make someone do, feel, or believe?
Most disappointing AI videos are not generation failures. They are planning failures. The shots look fine individually, the model behaved, the render completed — and the result is still forgettable because nobody decided in advance what the sequence was for.
A useful place to start is a single sentence, written before any prompt exists: After watching this, the viewer will ___. Not "understand our product" but "request a demo," "save the post," "stop using three separate tools," "feel uneasy about their current backup routine." If you cannot finish that sentence in under fifteen words, the video is not ready to be written, let alone generated.
From there, define the constraint set that shapes everything downstream: target length, aspect ratio, platform, whether there is narration or dialogue, whether a human appears on screen, and how much of the video depends on text overlays. A 9:16 clip for a vertical feed and a 16:9 explainer for a landing page are different genres with different pacing rules, even if they share footage.
The Four Layers of an AI-Ready Script
Traditional screenwriting moves from premise to outline to pages. AI production adds a fourth layer, because the script has to be readable by both a human editor and a generation model. Think of it as four stacked layers, each one more specific than the last.
Layer 1: Intent, audience, and the after-state
This is the strategic layer. Who is watching, what do they already believe, what objection are they carrying, and what is the smallest believable shift you can cause in sixty seconds? Write this down in plain prose. It never goes into a prompt, but every later decision references it.
Layer 2: Beat structure
Beats are emotional and informational units, not shots. A thirty-second piece usually has four to six. A three-minute explainer might have ten to fourteen. Beats are where you decide the order of revelation — what the viewer learns first, what you withhold, and where the turn happens.
Layer 3: Shot-level detail
Here the script becomes visual. Each beat breaks into one or more shots, and each shot gets a purpose, a subject, an action, a camera treatment, and a duration. This is the layer most people skip, and it is the layer that determines whether your generations come back usable.
Layer 4: Prompt-ready packaging
Finally, each shot is rewritten into language a video model can act on: subject first, then action, then camera, then light, then style, then technical parameters. Same information as Layer 3, different grammar.
Keeping these layers separate matters because they fail differently. A weak Layer 1 produces a technically polished video nobody cares about. A weak Layer 3 produces beautiful shots that do not cut together. A weak Layer 4 produces shots that ignore your instructions entirely.
Build a Beat Sheet Before You Write a Single Line of Dialogue
A beat sheet is a table with one row per beat and five columns: time range, beat name, what the viewer learns or feels, visual idea, audio idea. That is it. It is deliberately low-resolution so you can throw it away and rebuild it in ten minutes.
A reliable skeleton for short-form work:
- Hook (0–3s). The single most visually or verbally arresting thing you have. No logo, no throat-clearing.
- Promise (3–8s). Tell them what they will get if they stay. Specific beats vague.
- Tension (8–15s). The problem, the cost, the friction, the thing that is currently annoying them.
- Turn (15–22s). The moment the video earns its keep — a demonstration, a reveal, a counterintuitive claim.
- Proof (22–28s). One concrete piece of evidence: a number, a before-and-after, a short testimonial.
- Payoff + ask (28–30s). Resolve the tension and make one request. One.
For longer explainers, stretch the same shape: hook, context, problem, two or three proof beats, objection handling, demonstration, summary, ask. The proportion changes; the logic does not.
The beat sheet is also where you catch the most common structural flaw: a video that spends its opening explaining itself. If your first beat is a title card or a mission statement, you have written an introduction, not a hook.
Scene Blocking: The Bridge Between Story and Generation
Scene blocking is where a script stops being literature and becomes production instructions. Each scene block should answer a fixed set of questions so that nothing is discovered mid-generation.
- Scene ID and purpose. S03 — show the friction of manual work.
- Location and time of day. Small home office, early morning, window light from the left.
- Subject. Who or what is on screen, described identically every time they appear.
- Action. One dominant physical action per shot. "She closes the laptop and exhales" is one action. "She closes the laptop, stands, walks to the window, and answers a call" is four.
- Camera. Framing, movement, lens feel. Choose from a small vocabulary: wide static, medium push-in, over-the-shoulder, handheld follow, macro insert, aerial establish.
- Lighting and palette. Warm practicals, cool daylight, high-key studio, low-key with rim light.
- Duration. Plan in seconds, and plan short — most AI models degrade with long continuous motion.
- Audio. Narration line, dialogue line, ambient bed, or silence.
A concrete example. For a thirty-second productivity tool teaser, the tension beat might block as:
S03 | Medium shot, slight push-in | Home office, morning | Subject: same woman as S01, gray cardigan, hair tied back | Action: she scrolls a messy list, stops, rubs her temple | Camera: 50mm feel, gentle dolly in | Light: soft window light, warm highlights | Duration: 3s | Audio: ambient room tone, single low synth note enters
Notice how much of that is decisions, not description. Every line exists to remove a choice from the generation step. That is the whole point of blocking: spend your thinking here, where changes cost seconds, instead of in the edit, where changes cost hours.
Writing Prompts That Survive a Change of Model
Different video models respond to different phrasings. Some reward cinematic vocabulary and camera terminology; others respond better to literal, plain description. You cannot eliminate that variance, but you can write shot prompts that transfer well.
Order matters. Lead with the subject and their distinguishing features, then the action, then the camera, then the lighting, then the aesthetic, then the technical spec. Models weigh early tokens more heavily, so put the non-negotiables first.
One dominant action per prompt. Multi-step action is where limbs multiply and props teleport.
Describe visible behavior, not interior states. "Nervous" is unrenderable. "Checks the phone twice, taps a pen against the desk" is renderable. Convert every emotion into a physical tell.
Anchor identity once, reuse it verbatim. If your subject description is "woman in her thirties, gray cardigan, dark hair tied back, small scar above left eyebrow," that string should appear character-for-character in every shot she appears in. Paraphrasing is how characters drift.
Name the negative space sparingly. Long lists of what you do not want tend to introduce the very thing you are excluding. Keep it to two or three genuine dealbreakers.
Specify format explicitly. Aspect ratio, shot duration, and frame rate if the model supports it. Guessing wastes iterations.
Separate the what from the how. A useful habit is to write the shot description in two paragraphs: one about content, one about treatment. When a model ignores your style direction, you can swap the treatment paragraph without rewriting the content.
Continuity: The Hardest Part of Multi-Shot AI Video
Single-shot clips are easy. Sequences are where projects fall apart, because viewers track consistency ruthlessly — even when they cannot articulate what feels wrong.
Character continuity is the most visible failure. Build a character bible with one canonical description, two or three reference images, and fixed wardrobe. Lock the description string and paste it, never retype it.
Environmental continuity covers time of day, weather, and set dressing. If a scene is set at dusk in shot one and high noon in shot three with no narrative reason, the sequence reads as assembled rather than directed.
Screen direction is the rule people forget. If your subject moves left to right in one shot, they should generally keep moving left to right in the next. Crossing the line mid-sequence disorients the viewer and makes the cut feel like a mistake.
Color continuity can be handled after generation with a shared look: one LUT, one contrast curve, one saturation target applied across the whole piece. Cheaper than re-generating, and usually more convincing than trying to match color in prompts.
Prop continuity deserves its own checklist. The coffee cup, the notebook, the laptop, the jacket — write them into the scene block so they do not silently change between shots.
A practical trick is to plan sequences around fewer locations and fewer wardrobe changes than a live-action shoot would need. Constraints that would feel lazy in filmmaking read as coherence in AI video.
Sound Is Half the Script, Not an Afterthought
If audio is planned last, you will cut visuals to fit narration that was never designed for them. Write the audio column while you write the shots.
For narration, budget roughly two to two and a half words per second at a comfortable pace. A thirty-second video with a solid hook has room for about sixty to seventy words total, and that includes the ask. Write it, read it aloud with a timer, and cut a third.
Dialogue introduces lip-sync constraints. Keep spoken lines short — under ten words where possible — and frame speakers in ways that forgive small mismatches: medium shots, profile angles, hands near the face, cutaways on the stressed syllable.
Then plan the layers that carry emotion without words: room tone, footsteps, keyboard clicks, a single low drone under the tension beat, and deliberate silence before the turn. Silence is the cheapest dramatic device available and the most frequently skipped.
Finally, decide captions and on-screen text up front. Text is not decoration; it is a second narration track. Shots that hold text need more visual breathing room, which is a blocking decision, not a post-production fix.
Turn the Plan Into a Repeatable Pipeline
Once the structure is in place, the workflow becomes mechanical:
- Brief. One page: audience, after-state, length, aspect ratio, tone, constraints.
- Beat sheet. One table, throwaway quality, approved before writing.
- Scene blocks. One row per shot with all fields filled.
- Prompt pack. Each shot converted to a transferable prompt with locked identity strings.
- Generation. Two to four variants per shot, named by scene ID and version.
- Selects. Keep a keeper log noting which variant won and why — this becomes your reference for future projects.
- Assembly. Cut on the beat sheet timings first, then adjust for natural rhythm.
- Sound and text. Narration, SFX, music, captions, lower thirds.
- Consistency pass. Color match, audio levels, text placement, loudness normalization.
- Delivery. Export per platform and archive the project folder with prompts attached.
Naming conventions are the unglamorous hero here. S03_v02_medium_pushin tells you everything six weeks later. final_final2 tells you nothing. Version prompts alongside renders so you can reproduce a shot that worked.
Mistakes That Wreck Otherwise Good Scripts
Writing prose instead of shots. A paragraph of atmosphere is not a script. If you cannot point at a frame, it is not ready.
Too many ideas per shot. Density feels efficient and renders as mush. One idea, one shot.
No hook. Viewers decide in the first second or two. A slow build is a luxury of content people already chose to watch.
Ignoring the delivery format. Vertical framing changes composition, text size, and how much visual information survives. Plan for the screen it will actually appear on.
Character descriptions that drift. Paraphrase once and your lead becomes a different person by shot six.
Vague verbs. "Moves gracefully" gives a model nothing. "Turns her head toward the window" gives it a target.
Audio planned last. The result is a video that looks fine and sounds like an assembly.
Treating a bad shot as a script problem when it is a model problem. Sometimes you need a different take, not a different plan. Learn to tell the difference: if the shot reads correctly but renders poorly, regenerate; if the shot renders correctly and still feels wrong, revise the block.
Pre-Generation Quality Check
Run this list before you spend an afternoon generating:
- Can you state the after-state in one sentence?
- Does the hook land in the first three seconds without context?
- Does every beat have a visual idea, or are some just narration?
- Does every shot have one dominant action?
- Is the subject description copy-pasted identically everywhere?
- Are screen direction and time of day consistent?
- Is the narration under the word budget for the runtime?
- Does the video make exactly one ask?
- Could an editor cut this from the scene blocks alone, without asking you a question?
Any "no" is cheaper to fix now than after generation.
FAQ
How long should an AI-generated video be?
As short as the idea allows. Thirty to sixty seconds covers most social and product work. Longer pieces succeed when they carry genuine instructional content, not padded narrative.
Do I need a full screenplay before generating?
No. A beat sheet plus scene blocks is usually enough. Full screenplay formatting helps for dialogue-heavy narrative work, but it adds little to a product explainer.
Why do my characters change between shots?
Almost always because the identity description was rewritten rather than reused. Lock one description string, add reference images, and keep wardrobe fixed across the sequence.
Should I generate first and write later?
Occasionally useful for exploring a look. As a default method it produces sequences without structure, because you end up writing the story around whatever the model happened to give you.
How many variants per shot is reasonable?
Two to four. More than that usually means the prompt is too vague, not that the model is unlucky.
How do I handle shots that need text on screen?
Block them as their own shots with deliberate empty space, and write the exact text in the scene block. Retrofitting text onto busy frames is the most common cause of unreadable graphics.
What is the fastest way to improve?
Rebuild three existing videos using the beat sheet, scene block, and prompt pack structure. You will feel the difference in the first sequence.
A Seven-Day Practice Plan
Spend day one writing briefs and after-state sentences for five ideas — no visuals. Day two, build beat sheets for the two strongest. Day three, block every shot in one sequence, filling all fields. Day four, convert blocks into prompts and generate two variants each. Day five, assemble and add sound. Day six, run the consistency pass and ship it. Day seven, review the keeper log and note which instructions your chosen model honored and which it ignored.
That log is the real asset. Over a few projects it becomes a private manual for how your tools behave — and the scripts you write start anticipating it. That is the difference between using AI video and directing it.




