Commencer Gratuitement
Offre à durée limitée : forfaits annuels Starter et Basic à 50% de réduction 🎉

AI Storyboarding and Script Planning: A Director's Workflow

Oct 8, 2026

What a Virtual Director Actually Does

"Virtual director" sounds like a marketing phrase, but the idea underneath it is practical. Most of the work in a short film, a product ad, or an explainer video happens long before anything is rendered. Someone has to decide what the story is, how it moves, what the audience sees at each moment, and how every shot connects to the next. A language model paired with an image generator can carry a surprising amount of that load — not by replacing taste, but by removing the blank-page friction that stalls projects.

In practice, the role splits into four jobs:

  • Structural thinking. Turning a premise into a three-act shape, a beat sheet, or a scene-by-scene outline with clear emotional turns.
  • Shot planning. Converting script lines into a coverage list: which shots are needed, in what order, and why each one exists.
  • Visual reference. Producing frame-level stills that communicate framing, lighting, palette, and tone before production time is spent.
  • Continuity tracking. Remembering what a character wore in scene two so scene nine does not silently contradict it.

A human director still owns every final decision. The value is speed and iteration depth: you can test five endings in an afternoon instead of arguing about one for a week. That shift matters most for small teams, where one person is usually writing, planning, and reviewing at the same time.

Map the Pipeline Before You Generate Anything

The single most common failure with AI-assisted video is jumping straight to prompts. Someone types a beautiful-sounding sentence into an image tool, gets a striking frame, and then discovers the frame belongs to no coherent story. The fix is boring but effective: define stage gates and do not skip ahead.

A workable pipeline looks like this:

  1. Premise and logline. One sentence: who wants what, what blocks them, what it costs.
  2. Beat sheet. Eight to fifteen beats that carry the story from opening image to resolution.
  3. Script or narration draft. Dialogue, voice-over, or on-screen text, written to time.
  4. Shot list. One row per shot with framing, subject, action, duration, and purpose.
  5. Look board. Reference stills that lock palette, lighting direction, and texture.
  6. Animatic. Static frames cut to the real soundtrack so timing problems surface early.
  7. Generation. Moving shots produced from approved frames and prompts.
  8. Edit and sound. Assembly, grade, mix, captions, delivery specs.

Each gate has a review question. At the beat sheet stage: does every beat change something? At the shot list stage: does every shot earn its place, or is it decoration? At the look board stage: could a stranger describe the film's mood after five seconds? If the answer is no, fix it there. Fixing a beat costs a sentence. Fixing a generated sequence costs a day.

Write these gates down somewhere shared. A simple numbered document beats a chat history, because chat histories bury decisions that were made three days ago.

Stage One: From Logline to Beat Sheet

Structure is where models are genuinely strong. They have absorbed thousands of outlines and can propose shapes instantly, which is useful precisely because the first outline is never the right one.

Start with constraints rather than requests. Instead of "write me a story about a courier," try: "A 90-second sci-fi short. Single protagonist, one location, no dialogue. Two visual reveals, one at 60 seconds. Ending must reframe the opening image." Constraints force specificity, and specificity is what makes the output usable.

A useful beat sheet prompt sequence:

  • Ask for three competing structures. One linear, one in medias res, one structured around a repeating visual motif. Comparing options teaches you what you actually want.
  • Ask for the beats as a table. Columns: beat number, what happens, what changes, estimated duration. Tables expose pacing problems that prose hides.
  • Interrogate the weakest beat. Ask which beat is doing the least work and why. Models are often better critics than generators when asked directly.
  • Rebuild with your own ending. Replace the final two beats with your version and ask what breaks upstream. This is where you catch a twist that was never set up.

Duration estimates matter more than they seem. A 60-second piece typically supports five to eight story beats, not fifteen. If your outline implies twenty beats, you are writing a three-minute film and you should admit it now.

Stage Two: Script Drafting Without Losing Your Voice

The fastest way to make AI-written scripts feel generic is to ask for "natural dialogue." Natural to whom? You need to define voice as a set of observable rules.

Collect a voice brief before drafting:

  • Sentence length. Short and clipped, or long and winding? Give a target range.
  • Vocabulary register. Technical, conversational, formal, slangy? List five words the character would never say.
  • Rhythm. Does the character interrupt, trail off, or finish other people's sentences?
  • Subtext. What does the character want but refuse to name?

With that brief, drafting becomes a controlled process. Generate three versions of the same scene under three different rhythms, then steal the best lines from each. That hybrid approach produces something that reads like a writer's draft rather than a machine's summary.

Narration is easier than dialogue and usually stronger in AI-generated video, because lip-sync and performance nuance are hard to fake. If your piece can work as voice-over over strong visuals, it will be more convincing and far cheaper to iterate.

Time your draft out loud. Scripts run shorter than they look. A page of dialogue that reads as 45 seconds often lands closer to 70 once performed with pauses and breath. For narration, 140 to 150 spoken words per minute is a reasonable planning number.

Stage Three: Shot Lists and Coverage Planning

A shot list is the bridge between writing and generating. It is also the document most people skip, which is why their edits feel random.

Build it as a table with these columns: shot number, scene, subject, action, framing, camera movement, lens feel, duration, and purpose. The purpose column is the one that saves you. If you cannot write "establishes isolation" or "reveals the second coffee cup," the shot is probably decorative.

Coverage categories worth planning deliberately:

Category Typical use Notes
Establishing wide Place the scene geographically One per location, usually
Medium two-shot Relationship and blocking Carries most dialogue scenes
Close-up Emotion, detail, decision Use sparingly for impact
Insert Objects, hands, screens Cheap to generate, high value in edits
Transition shot Movement that bridges scenes Doorways, vehicles, light changes
Reaction shot Listener's face Keeps dialogue scenes alive

For AI production, favor inserts, wides, and slow movements. These categories hide generation artifacts well because nothing in frame demands anatomical precision for long. Fast action, crowded scenes, and complex hand interaction remain the hardest cases; plan around them rather than fighting them.

Shoot ratio works differently here. In live action you might shoot ten times your final runtime. In AI production, budget several generated attempts per usable shot, and keep your shot count modest — 15 to 25 shots for a 60-second piece is a comfortable target.

Stage Four: Storyboards, Look Development, and Visual Consistency

Storyboarding with image generation is fast, but only if your prompts carry the same information a storyboard artist would draw.

Use a repeatable prompt formula:

subject + action + wardrobe + framing + lens feel + lighting direction + palette + texture + mood

A blank version: middle-aged courier in a patched canvas jacket, stepping through a rain-slick alley door, medium shot, slight wide-angle, cold key light from the left, teal and amber palette, film grain, tense.

Two things make this work. First, the order is stable, so you can swap one element at a time and see exactly what changed. Second, nothing is vague — "cold key light from the left" produces more controllable results than "dramatic lighting."

Consistency is where most projects break. Practical techniques:

  • Character sheet first. Generate one clean reference for each main character in neutral light, front and three-quarter view. Reuse that description verbatim in every later prompt.
  • Lock wardrobe as text. Write it once — "mustard raincoat, scuffed white sneakers" — and paste it into every shot featuring that character.
  • Separate look from action. Generate the look with a simple action, approve it, then change only the action for subsequent shots.
  • Keep a location bible. Every location gets one approved wide and one detail shot. New angles are generated from those references, not from scratch.
  • Use image-to-image for new angles. Starting from an approved frame preserves lighting direction and color far better than a fresh text prompt.

Treat every approved frame as an asset, not a stepping stone. A frame you approve is a decision you should not have to make again.

Prompting Camera Language and Choosing the Right Model

Camera language translates well into text if you use the vocabulary of a real shot list. Instead of "cinematic," specify: slow dolly in, eye level, 40mm equivalent, shallow depth of field, subject centered then drifting left.

Useful movement phrases and what they signal:

  • Slow push in: growing tension, realization, intimacy.
  • Pull back: isolation, context reveal, ending.
  • Lateral tracking: momentum, following a journey, parallel action.
  • Handheld drift: unease, documentary immediacy.
  • Static locked-off: control, formality, comedy timing.

Model choice should follow the shot, not the other way around. Evaluate candidates against these criteria:

  1. Motion coherence. Does movement hold together across several seconds, or does geometry melt?
  2. Style fidelity. Can it hold your palette and texture across many shots, or does each generation drift?
  3. Image-to-video strength. If you have approved frames, how closely does the output respect them?
  4. Length per generation. Longer clips reduce stitching work but often lose detail.
  5. Iteration speed. A slightly weaker model that responds in seconds often beats a better one that takes minutes, because you will run it twenty times.
  6. Controllability. Does it accept camera, seed, or reference inputs that you can actually steer?

Assign models per shot type rather than choosing one for the whole project. A model that renders landscapes beautifully may fail at faces; a model that nails faces may produce flat, lifeless motion. Document which model produced which approved shot, because re-generating later requires the same tool and settings.

A Worked Example: Sixty Seconds, Eight Shots

Suppose you are making a 60-second short about a night-shift lighthouse keeper who receives a signal that should not exist.

Beats. Opening image of the light sweeping; routine established; anomaly detected; doubt; verification attempt; confirmation; decision; final image that mirrors the opening with a change.

Script. Narration only, roughly 140 words. Draft three versions — one procedural, one wistful, one clipped and technical. Blend: procedural opening, wistful middle, technical ending.

Shot list.

  1. Extreme wide, lighthouse on a black coastline, static, 6 seconds — establishes isolation.
  2. Medium, keeper at the console, slow push, 5 seconds — introduces the person.
  3. Insert, a dial needle jumping, static macro, 3 seconds — the anomaly.
  4. Close-up, keeper's face, handheld micro-drift, 4 seconds — doubt.
  5. Wide, empty sea through the window, static, 5 seconds — anticipation.
  6. Insert, logbook handwriting, slow tilt, 4 seconds — verification.
  7. Medium, keeper reaching for the switch, lateral drift, 5 seconds — decision.
  8. Extreme wide, the beam sweeping the same shot as shot one, now with a second light answering, static, 8 seconds — payoff.

Look board. One approved frame per shot, generated from a locked palette: deep navy, sodium orange, wet stone texture. Character sheet for the keeper: grey wool sweater, weathered hands, no visible face detail beyond silhouette and jaw.

Animatic. Cut stills to the narration. At this stage, shot eight almost always feels too short. Extend it to 10 seconds and trim shot five. You just saved a generation cycle.

Generation. Produce shots 1 through 7 with image-to-video from approved frames. Generate shot 8 as a variant of shot 1 with one prompt change. Multiple attempts each, keeping every version until the edit is locked.

Quality Control, Versioning, and Common Mistakes

Once you are generating dozens of clips, organization becomes the difference between a finished film and a folder of near-misses.

A naming convention that survives contact with reality:

project_shot03_v04_model_approved.mp4

Keep three buckets: work, approved, final. Nothing moves to approved until it passes a checklist:

  • Does the shot match the approved look frame?
  • Any warping, extra fingers, or melting edges?
  • Does the camera move match the shot list?
  • Does it cut cleanly against the shots on either side?
  • Is the duration within a second of the plan?

Common mistakes, in rough order of how much time they waste:

  1. Generating before the outline is locked. Rework cascades through every downstream asset.
  2. Describing mood instead of image. "Epic" is not a prompt; "low horizon, backlit silhouette, orange gradient sky" is.
  3. Changing five prompt variables at once. You lose the ability to learn what worked.
  4. Ignoring duration. Beautiful 8-second shots do not fit a piece edited to 4-second rhythm.
  5. Using one model for everything. Specialize per shot type.
  6. Skipping the animatic. Timing problems found in the edit are the most expensive kind.
  7. No continuity document. Characters quietly change jackets and eye color.
  8. Overloading the frame. Complex crowds and rapid action remain unreliable; design around limits.
  9. Deleting failed generations too early. A rejected clip often becomes the fix for a different shot.
  10. No sound pass. AI video is unforgiving without atmosphere, foley, and music; silence makes even good shots feel fake.

Finally, plan a delivery pass: consistent resolution and frame rate, a single grade across all shots, captions if anyone might watch muted, and a loudness-normalized mix. Technical polish is what makes an AI-assisted piece read as intentional rather than experimental.

FAQ

Do I still need a script if I am generating everything?
Yes, perhaps more than ever. Generated clips have no memory of intent, so the script and shot list are the only things holding the piece together. A written beat sheet also makes it obvious when a shot exists for no reason.

How many shots should a 60-second video have?
Fifteen to twenty-five for a visually dense piece, eight to twelve for something slow and atmospheric. Fewer shots with longer holds are easier to keep consistent, so start small and add coverage only where the edit feels thin.

Can I keep the same character across many shots?
Reasonably well, with discipline. Lock a written wardrobe and appearance description, generate one approved reference, and build every new angle from that reference with image-to-image rather than a fresh prompt. Expect to regenerate some shots regardless.

Which is more important, the model or the prompt?
The prompt, by a wide margin, at least until you reach the top tier of tools. A precise prompt on an average model beats a vague prompt on a great one almost every time, because precision is what makes results repeatable.

How long does this workflow take in practice?
For a 60-second piece, budget one day for outline and script, half a day for the shot list and look board, one to two days for generation and retries, and one day for edit, sound, and polish. Most of the variance lives in retries and in how early you lock the outline.

What should I learn first?
Shot language. Understanding framing, lens choice, and movement does more for generated video quality than any prompt trick, because it tells you what to ask for and how to judge whether you got it.

Alexander

Alexander