Why AI Belongs in Pre-Production, Not Just Post
Generative video tools get most of their press for spectacle: impossible camera moves, dreamlike landscapes, characters who morph into something else entirely. That is the demo reel. The working reality is quieter and far more useful. The highest-leverage place to use AI in video production is not the final render or the color pass, it is pre-production, where decisions are cheap and reversible.
A weak logline costs nothing to change. A weak finished video costs you days, a budget, and a publishing slot that you cannot get back. AI is very good at helping you make those cheap decisions faster: expanding a rough idea into five structured directions, converting a script into a shot list with camera, lens, and lighting notes, generating storyboard frames for review before anyone commits to a shoot or a generation queue.
There is one job AI does badly, and it is worth naming early: deciding what the video is actually about. Models are excellent pattern completers. They will happily produce a polished, generic script about productivity, coffee, or the future of work. They will not tell you that your specific audience cares about one narrow, uncomfortable, specific thing. That part is still on you.
The workflow in this guide treats AI as a drafting partner with unlimited stamina and no taste of its own. You provide the taste.
Define the Attention Contract Before You Write Anything
Before a single prompt, write down five constraints. Every one of them will make your AI drafts dramatically better, because it removes the ambiguity that models fill with clichés.
- Audience. Not demographics, but a situation. A first-time freelancer pricing a project. A marketing lead who has already tried three tools.
- Platform and aspect ratio. Vertical short-form and horizontal explainer are different crafts. The frame shape changes what can be communicated in a single shot.
- Runtime. A 30-second script has roughly 70 to 90 spoken words. A 60-second script has about 150. Knowing the number prevents the classic AI failure of writing three minutes of material for a 45-second slot.
- The promise. One sentence that says what the viewer gets by staying. If you cannot write it, the video does not have one.
- The emotional register. Calm authority, dry humor, urgent curiosity, warm reassurance. This is the single most under-specified input in most AI video projects, and it is the one that most changes the output.
Write these five lines in a plain text file and keep them open. Then paste them into the top of every AI conversation you have about this video. Call it the brief block. It is the cheapest quality improvement available to you.
From Idea to Story Skeleton in Four Passes
Trying to get a finished script from one prompt is the most common beginner mistake. You get something that reads smoothly and says nothing memorable. Instead, run four passes, each with a different ask.
Pass 1: Logline and Angle Proposals
Ask the model for ten one-sentence loglines for the same topic, each approaching it from a different angle: a contrarian take, a personal failure story, a myth-busting take, a step-by-step take, a cost-of-inaction take. Do not ask for the best one. Ask for range, then choose. Ten options take a minute to generate and give you a genuinely strategic choice rather than a stylistic one.
Pass 2: The Beat Sheet
Once you have a logline, ask for a beat sheet of five to seven beats. For a short video, that is typically hook, setup, first value point, second value point, proof or example, payoff, and a closing call to action. For a longer piece, beats become acts. The critical instruction here is to make the hook a specific claim or a specific tension, not a greeting. Never open with a question the audience can answer with obviously yes or no.
Pass 3: Scene Cards
Expand each beat into a scene card of three to five lines: where we are, what the viewer sees, what the viewer hears, and what changes by the end of the scene. Scene cards are the single most useful artifact in the whole process, because they are human-readable and machine-readable at the same time. You can hand them to an editor, a generator, or another AI tool.
Pass 4: Script Draft
Only now write the actual script. Feed the scene cards back and ask for a draft in the emotional register you defined, with a target word count, and with an explicit instruction to avoid filler openings such as an introduction or a restatement of the title. Ask for two versions with different pacing: one conversational, one clipped and punchy. Compare them aloud.
Reading the Script Out Loud Is a Technical Step
This is not a stylistic flourish. Reading the draft aloud is a diagnostic. It surfaces sentences that are grammatically fine and physically unsayable. It reveals where a breath falls in an awkward place, where a list of three items becomes a list of seven, and where a joke that looked sharp on the page dies in the mouth.
Mark every sentence you stumble over. Ask the model to rewrite only those sentences, keeping everything else untouched. Iterating on the whole script when you only dislike four sentences is how drafts drift away from the original idea.
A practical rhythm: read aloud, mark, rewrite marked lines, read again. Two cycles is usually enough. Three cycles means the underlying structure is wrong, and you should return to the beat sheet rather than polishing prose.
Translating a Script Into a Shot List
A script says what happens. A shot list says what the camera does. The gap between them is where most AI-assisted videos become visually incoherent, a sequence of unrelated beautiful images rather than a story.
Build your shot list in columns, whether in a spreadsheet or a plain document: shot number, script line, shot size, camera movement, subject and action, lighting, duration, and asset source. Eight columns sounds heavy. It takes about twenty minutes for a short video and saves hours of regeneration later.
Two rules make shot lists work. First, every shot must change something: new information, new location, new emotional temperature. If two adjacent shots show the same subject in the same framing doing the same thing, delete one. Second, assign a duration budget before generation. AI video tools tend to produce clips in fixed increments, and a coherent edit needs variation in shot length, roughly a mix of one-second inserts, two-to-three-second beats, and occasional longer holds.
Using AI to Draft the Shot List
Paste the script and ask for a shot list in those columns, with an instruction to alternate shot sizes and to include at least one insert shot per beat. The first output will be competent and slightly boring. Then ask a sharper question: where would a cut feel jarring, and what would justify it? Models are surprisingly good at identifying pacing risk when you ask directly.
Storyboarding Without Drawing Skills
A storyboard does not need to be beautiful. It needs to answer one question fast: does this sequence read? You can generate storyboard frames with an image model in a fraction of the time it takes to sketch them, provided you control style.
Use a style lock: a short, repeated descriptor appended to every frame prompt that fixes medium, palette, contrast, and lens character. Something like a muted three-color palette, soft overcast light, 35mm framing, photoreal but slightly desaturated. Consistency in the board matters more than beauty, because the board is a communication tool.
For character continuity across frames, describe the character in the same words every time and keep a reference image in the conversation. Models drift toward generic faces otherwise, and a board with three different-looking protagonists is worse than a board with stick figures.
Anatomy of a Shot Prompt That Actually Works
Most failed generations come from prompts that describe a mood and nothing else. A usable shot prompt has six parts, and you can assemble them mechanically.
- Subject and action. Who or what, doing what, in the present tense. Specific verbs beat adjectives every time.
- Setting and time. Interior or exterior, era, weather, time of day, and one concrete environmental detail.
- Framing and lens. Wide, medium, close, macro; shallow depth of field, wide-angle distortion, telephoto compression.
- Camera behavior. Static, slow push in, handheld follow, orbit, crane down. Choose one and only one primary move.
- Lighting and palette. Direction of light, hardness or softness, dominant and accent colors.
- Continuity anchors. Wardrobe, props, hair, and color notes that must persist across shots.
Then add constraints separately: what must not appear, what must not move, and what must remain stable across the clip. Negative constraints are not a substitute for a clear positive description, but they prevent specific, recurring failures.
The Iteration Loop That Saves Time
Change one variable per generation. If you alter the framing and the lighting and the action simultaneously, you will not know which change broke the shot. Keep a running log of prompts and results for the shots that matter. Twenty lines of notes will teach you more about a given model than a week of watching other people's outputs.
Matching the Model to the Moment
Different video models are genuinely good at different things, and the differences are practical, not tribal. Rather than chasing the newest release, classify your shots and match them.
- Photoreal human performance. Prioritize models with strong face stability and natural micro-movement. Test with a talking-head clip before committing a whole scene.
- Stylized or illustrated worlds. Prioritize models with strong style adherence and clean edges. These often handle bold graphic looks better than photoreal ones.
- Complex camera moves. Prioritize models with reliable motion coherence across frames, especially for orbital and tracking shots.
- Short inserts and texture shots. Almost any model handles these well, so use your fastest, cheapest option. Do not spend your best tool on a three-frame coffee pour.
- Long continuous takes. Treat with suspicion across the board. If you need more than a few seconds of unbroken action, generate in segments and cut.
Once you have classified, build a default pipeline: one model for hero shots, one for inserts, one for stylized sequences. Consistency of pipeline produces consistency of look, which is worth more than squeezing out marginal quality from a different tool every scene.
The Edit Is Where AI Video Usually Falls Apart
Generation is not editing. A folder of strong clips can still become a weak video, and it usually does for the same three reasons.
Pacing without rhythm. AI clips often share a similar internal tempo: smooth, slow, evenly weighted. Cut against that. Insert a hard cut on a beat, let one shot run long, then cut twice quickly. Rhythm comes from contrast, not from consistent smoothness.
Silence that reads as emptiness. Sound design carries more perceived production value than image quality. Add room tone, a subtle low-frequency bed, and tactile foley for anything the audience sees touching something. If you are using a voiceover, cut the music under the first two seconds of every important line.
Captions as an afterthought. A large share of viewers watch with sound off. Burn in or style captions so they match the visual system: same font family as any on-screen text, high contrast, safe-area aware. Automatic captioning is a starting point, not a final pass. Read every line for punctuation and line breaks.
Seven Failure Modes and Their Fixes
- Morphing anatomy. Fix by shortening the clip, simplifying the action, and framing closer so hands are out of frame.
- Flicker and texture crawl. Fix by reducing high-frequency detail in the prompt and avoiding busy patterned backgrounds.
- Wardrobe and hair drift. Fix with explicit continuity anchors repeated in every prompt for the scene.
- Ignored prompt elements. Fix by reducing the prompt to the three most important ideas, then adding back one element at a time.
- Unmotivated camera movement. Fix by choosing a single move that a human operator would actually make for a reason.
- Flat lighting. Fix by naming the light source and its direction rather than asking for good lighting.
- Generic performance. Fix by describing a specific micro-action instead of an emotion: not nervous, but tapping a pen twice against a closed notebook.
A Repeatable Weekly Workflow
A sustainable rhythm beats a heroic sprint. A working pattern looks like this: brief block and loglines on day one; beat sheet and scene cards on day two; script draft, read-aloud pass, and shot list on day three; storyboard frames and style lock on day four; generation in batches on day five; edit, sound, and captions on day six; publish and log results on day seven.
The logging step is the one everyone skips and the one that compounds. Record the hook style, the runtime, the platform, and the retention you observed. After ten videos you will have a private dataset that no generic prompt can beat.
FAQ
Can AI write a complete video script I can use as-is? It can produce a complete draft in seconds. Treat it as a first draft from a fast, well-read, slightly generic writer. The structural work, the read-aloud pass, and the specificity are what turn it into something worth publishing.
How long should an AI-generated video clip be? As short as the story allows. Most clips in a finished edit run between one and four seconds. Generate longer clips only when the action genuinely needs to breathe, and expect to trim.
Do I still need a storyboard if I am generating everything? Yes, and arguably more so. Generation is fragmented by nature, and the board is the only artifact that shows you whether the fragments form a sequence before you spend time producing them.
How do I keep characters consistent across shots? Lock a written description, reuse a reference image, and repeat wardrobe and hair details in every prompt. Also accept that minor variation is normal in generation, so choose shots where small differences are not distracting.
What is the fastest way to improve output quality? Improve the input. A specific action, a named light direction, and one clear camera move will outperform any stack of quality-boosting adjectives.
Should the voiceover be AI or human? For short, high-stakes, brand-facing pieces, human performance still carries emotional range that is hard to fake. For rapid tests, internal explainers, or heavily processed styles, synthetic voice is efficient and often indistinguishable.
The through-line in all of it is unglamorous: write the brief, structure the story, plan the shots, then let AI do the heavy lifting it is actually good at. Tools change every few months. That sequence does not.


