Start Free Now
Limited Time Offer: Get 50% OFF Starter & Basic Yearly Plans 🎉

From Concept to Clip: AI Video Story and Shot Design Workflow

Sep 23, 2026

Why the gap between story and shot is where AI video projects fail

Generative video tools have become genuinely impressive. Type a sentence, wait a minute, and you get a moving image with believable light, plausible physics, and a camera move that would have cost a fortune a decade ago. And yet most AI video projects still collapse somewhere between the idea in your head and the clip on your timeline. The failure almost never happens at the generation step. It happens in translation.

A story idea is abstract. A shot is concrete. Between those two things sit dozens of decisions: who is in frame, how close the camera is, what the light is doing, how long the moment lasts, what came immediately before it, and what has to match it three shots later. When you skip those decisions and hand a vague paragraph to a video model, the model makes all of them for you — quickly, confidently, and inconsistently. Shot one gives you a wide sunlit street. Shot two gives you the same street at a different time of day with a different-looking character. Shot three introduces a third visual language entirely.

The practical fix is not a better prompt. It is a pipeline. Treat the work as four translation layers — story, sequence, shot, render — and make each layer produce a written artefact the next layer can consume. This guide walks through that pipeline in detail, with prompt structures, decision criteria, worked examples, continuity systems, and the mistakes that cost the most time.

The four-layer pipeline: story, sequence, shot, render

Think of the pipeline as a funnel that gets more specific as it narrows. Each layer has a single job, a clear output, and a natural review moment.

Layer 1: Lock the story spine

Before any visual thinking, write the spine in plain language: a protagonist, a want, an obstacle, a turn, a resolution. Ten to fifteen sentences is enough for a short film or a commercial. The spine is the reference document you return to whenever a shot feels beautiful but pointless. If a shot cannot be traced back to a line in the spine, cut it.

Layer 2: Map sequences to time and place

Break the spine into three to six sequences. A sequence is a unit of continuous time and place — a morning routine, a negotiation, a commute. For each sequence, note the location, the time of day, the emotional temperature, and the single thing that must change by the end. This is where you decide visual rules: is the film warm and handheld, or cold and locked off? Write those rules down. They become your global prompt modifiers.

Layer 3: Write shot cards

Each sequence becomes three to eight shots. A shot card is a compact structured description: subject, action, shot size, camera move, lens feel, lighting, duration, and continuity notes. Shot cards are the real product of pre-production. They are also the only thing you should ever paste into a video model.

Layer 4: Render in passes, not in one prompt

Attempting to generate a finished, edited, graded clip from a single prompt is the most common beginner error. Instead, render in passes: a blocking pass to check composition and motion, a detail pass to refine performance and light, and a finishing pass for polish. Each pass narrows the search space.

Where the review moments live

Put a review gate between every layer. A fifteen-minute review of the sequence map saves hours of re-rendering later. Reviews are cheap on paper and expensive on the render queue.

How to write a shot card a video model can actually follow

A shot card is not poetry. It is a specification. The best ones read like a note a cinematographer would nod at: specific, ordered, and free of ambiguity.

The six required fields

  1. Subject — who or what occupies the frame, plus one identifying detail (wardrobe, colour, an object they carry).
  2. Action — one primary verb, one secondary beat. "She opens the envelope, then looks up" is directable. "She feels uncertain" is not.
  3. Shot size — wide, medium, close, extreme close. Pick one.
  4. Camera — static, slow push, tracking, handheld drift, crane, orbit. Pick one and include a speed qualifier.
  5. Light and time — soft window light, golden hour backlight, overcast diffusion, practical neon.
  6. Duration — in seconds, and whether the motion resolves inside that window.

The optional fields that punch above their weight

Lens character (shallow depth of field, wide-angle distortion, long-lens compression), colour palette, weather, foreground occlusion, and screen direction. Foreground occlusion in particular — a doorway edge, a passing car, a shoulder — adds production value that models produce reliably when asked.

A worked micro-example

Weak: A woman walks through a market and thinks about her decision.

Strong: Medium shot, static, slight handheld drift. A woman in a mustard jacket walks left to right through a busy morning market, holding a folded paper. Stalls blur in the foreground. Overcast diffused light, cool palette with warm accents. 4 seconds. She slows, then stops mid-frame.

The second version gives the model a subject, an action, a framing, a movement, a light condition, a duration, and a screen direction. It is also short enough to survive model token limits without losing the important clauses.

Choosing the right model and mode for each shot

There is no single best video model, only a best match for a shot. Build a simple decision table and stop re-litigating the choice every time.

Text-to-video vs image-to-video vs video-to-video

Text-to-video is fastest for exploring blocking and mood, but the least controllable for recurring characters. Image-to-video is the workhorse for character consistency: generate or shoot a still first, approve it, then animate it. Video-to-video is for restyling, frame-rate changes, and matching an existing plate — useful when you already have a locked edit and want to change the look without changing the timing.

A reliable default for narrative work is: lock a still per shot, animate it, then reserve text-to-video for inserts, textures, and abstract transitions.

Resolution, duration and aspect ratio decisions

Generate at the highest resolution your time budget allows for hero shots, and at lower resolution for anything that will be small on screen or buried under an effect. Keep a single aspect ratio across the project unless the format demands otherwise; mixed ratios create ugly framing surprises in the edit. Keep individual generations short — three to six seconds — and stitch. Long single generations drift, morph, and lose the subject.

When to combine tools

Mix models on purpose. One tool may handle faces and skin better, another may excel at landscapes and camera movement, a third at stylised animation. Write the choice into each shot card so the pipeline stays repeatable and you can explain a shot's provenance months later.

Budgeting generation time proactively

Assume roughly a third of your generations will be unusable. If a sequence needs six usable shots, plan to produce eighteen. Queues are slow at peak hours, so batch similar shots together and let them render overnight while you work on sound design.

Continuity: the discipline that separates amateur from professional

Continuity is the invisible work that makes AI video feel authored rather than assembled. It is also the cheapest thing to fix on paper and the most expensive to fix after rendering.

Character continuity

Create a character sheet: face reference, wardrobe, hair, one signature object. Reuse the same reference image across every shot featuring that character, and keep the description text identical — word for word. Paraphrasing your own character description is one of the fastest ways to produce a stranger in the next shot.

Location and prop continuity

Write a location sheet with the same discipline: architectural style, wall colours, key furniture, weather, time of day. Props that cross shots need their own note — the red mug in shot two must be the red mug in shot nine. Track screen direction too: if a character exits frame right, they should enter frame left in the next shot of the same scene.

Lighting and grade continuity

Decide the film's lighting logic up front. Is sunlight always coming from the window side? Does the palette shift warmer as the story resolves? A short grade pass at the end, applied across all shots, hides a surprising number of small inconsistencies.

Build a continuity ledger

Keep a simple table: shot ID, character, wardrobe, location, time of day, palette, key props. Twenty rows of this will save you a full day of re-rendering. It also makes it possible to hand the project to a collaborator without a two-hour briefing.

Camera language: a practical vocabulary and prompt patterns

Models respond well to a small, disciplined vocabulary. Learn twelve terms and use them consistently.

Shot sizes

Extreme wide, wide, full, medium, medium close-up, close-up, extreme close-up. If you only specify two, specify medium and close-up — they carry most dialogue and emotion.

Moves

Static (locked off), slow push in, slow pull out, lateral tracking, follow, handheld drift, crane up, orbit, rack focus. Always add a rate: very slow, deliberate, snap. Rate is what makes a move feel intentional rather than accidental.

Lens and depth

Shallow depth of field isolates the subject and reads as premium. Deep focus reads as documentary. Wide-angle distortion suggests unease or comedy. Long-lens compression flattens distance and can make a crowded street feel intimate.

Lighting descriptors that actually work

Soft window light, hard midday sun, golden hour backlight, overcast diffusion, practical neon, single-source low key, bounced fluorescent. Pair each with a colour temperature word: warm, neutral, cool, mixed.

Two reusable prompt templates

Narrative shot: [Shot size], [camera move + rate]. [Subject with one identifying detail] [primary action], then [secondary beat]. [Location and time of day]. [Lighting]. [Lens character and palette]. [Duration].

Transition/insert: [Shot size], [camera move]. [Object or texture detail] in [location]. [Lighting]. [Duration]. No people.

Keep the order stable. Consistency in structure trains you to spot what is missing.

Worked example: a 45-second launch film from one paragraph

Start with a paragraph: A cyclist leaves the city before dawn to deliver a prototype to a workshop in the hills. On the way, the road opens up and the ride becomes joyful. She arrives as the sun rises, hands over the package, and the team unboxes it.

Spine. Eight sentences. Want: deliver the prototype. Obstacle: distance and time. Turn: the road opens. Resolution: sunlight, the handoff, the unboxing.

Sequence map. Three sequences: pre-dawn city, open road at first light, sunrise workshop. Visual rules: cool blue-to-amber progression, shallow depth of field, slow deliberate camera moves, no handheld after sequence one.

Shot cards. Sequence one: four shots — a close-up of hands on handlebars, a wide of empty streets, a medium tracking shot, a rear-view of her leaving frame. Sequence two: five shots — a wide valley reveal, a close-up of a spinning wheel, a low tracking shot, a silhouette against the rising sun, an overhead of the road curve. Sequence three: four shots — a wide of the workshop, a medium of the handoff, a close-up of the box opening, a wide of the team around the table.

Rendering passes. Blocking pass for all thirteen shots at low resolution. Approve composition and motion. Detail pass at full resolution for the six hero shots. Finishing pass for grade and grain.

Edit. Cut on action and match the wheel close-up to the rhythm of the music. Keep the pre-dawn sequence cold and slow, let sequence two open up, and end on a held wide.

The whole film is thirteen shots, roughly six minutes of usable screen time, and about two hours of generation. Planned this way, it is manageable. Improvised, it becomes a weekend of guessing.

Review loops, versioning and asset hygiene

Discipline in file handling is what makes a project survivable after the first week.

Naming conventions

Use a consistent pattern: project_sequence_shot_version. For example, launch_s02_sh04_v03. Add a status tag — blocking, approved, rejected — either in the filename or a tracking sheet. Never overwrite an approved render.

A three-pass review cadence

Review at shot level for framing and motion. Review at sequence level for rhythm and continuity. Review at film level for pacing and story. Each pass answers a different question; mixing them into one marathon session produces muddled notes.

Approval gates

Do not move a shot into a finishing pass until its composition is locked. Do not lock the edit until all hero shots are approved. Gates feel bureaucratic on a small project and save days on a large one.

Keeping prompts with the assets

Store the exact prompt and settings alongside each rendered file. When a client asks for "the same shot but warmer", you will be able to reproduce it in one generation instead of twenty.

Common mistakes, and how to avoid them

Writing prompts instead of shot cards. A prompt is a sentence. A shot card is a specification. Always work from the latter.

Generating too long. Anything past eight seconds drifts. Generate short, cut on action.

Changing the character description between shots. Copy and paste the exact same text every time.

Mixing lighting logic mid-scene. Write the lighting rule per sequence and obey it.

Falling in love with an off-brief shot. Beautiful footage that does not serve the spine costs you the edit.

Rendering before locking blocking. Low-resolution approval first. Always.

Ignoring screen direction. Eye-line and exit-direction errors are the fastest way to make an edit feel broken.

Skipping the sound plan. Decide where the music swells and where sound drops out before you finish picture. It changes which shots you need.

No version naming. Untracked files become a folder full of mysteries within days.

Treating the model as the author. The model renders; you decide. Every shot card is a decision you made, and that is what makes the result yours.

FAQ

How many shots do I need for a one-minute video? Twelve to twenty for a paced narrative, fewer for a single continuous idea. Shot cards make the count obvious before you render anything.

Can I get consistent characters without a reference image? Sometimes, but image-to-video with a locked still is far more reliable. Generate the still, approve it, then animate.

What is the fastest way to test a visual style? Render three stills from three shots in the same sequence at low quality. If they feel like the same film, commit to the rule set.

Do I need editing software? Yes. Generation produces clips; editing produces films. Any timeline editor with basic colour tools is enough.

How long does a planned project take? A one-minute film with fifteen shots is typically two to four hours of generation plus two to three hours of editing, assuming shot cards are written first.

What if the model ignores part of my shot card? Cut the card down to its six required fields, re-render, and add complexity back one field at a time until you find the clause that broke it.

Should I regenerate or fix in the edit? Fix in the edit for framing, rhythm, and colour. Regenerate only for structural problems like missing action or wrong subject.

How do I keep a long project organised? One folder per sequence, one file per shot version, one row per shot in a continuity ledger. Simple systems survive; elaborate ones get abandoned.

Plan the story, map the sequences, write the shot cards, then let the models do what they are good at. The tooling will keep improving, but the translation discipline is what turns a concept into a clip that actually says what you meant.

Alexander

Alexander