Limited Time Offer: Get 50% OFF your first month of Pro & Ultra plans 🎉

Text-to-Video Production: Shot Lists, Models, Workflow

Sep 13, 2026

Why Text-to-Video Stopped Being a Novelty

A prompt box that turns a sentence into moving footage used to be a party trick. You typed something poetic, waited four minutes, and got six seconds of a half-melted face drifting through an impossible hallway. Everyone nodded politely and went back to their timelines.

That era is over. Modern text-to-video systems can hold a character's face across a cut, keep a camera move physically plausible for eight seconds, and follow a shot list you wrote in plain prose. The bottleneck has moved. It is no longer "can the model make video?" It is "can you direct the model like a director instead of a slot-machine player?"

This guide is about that second question. It covers how to choose a model for a shot rather than for a mood, how to write prompts that behave like shot lists, how to keep visual consistency across a sequence, and how to judge output before it reaches an editor's timeline. It is written for creators who already know what a close-up is and want output they can actually cut.

The Three Jobs Text-to-Video Models Actually Do

Before comparing tools, separate the work. Almost every clip you generate falls into one of three buckets, and each bucket rewards a different model behavior.

Shot generation. A single continuous moment: a woman opens a letter, a drone clears a ridge line, steam rises off a bowl. The deliverable is one usable take. These runs reward models with strong prompt adherence and clean motion physics.

Sequence generation. Four to ten related shots that share a look, a wardrobe, a location, or a character. The deliverable is a scene. These runs reward consistency features — image conditioning, reference frames, seed control — far more than raw resolution.

Motion transfer and animation. You already have a still image, a photograph, or a rendered frame, and you want it to move. The deliverable is animation applied to an existing asset. Here, image-to-video pipelines and motion-driven models win outright.

Most disappointing workflows come from using a sequence tool for shot work, or a shot tool for sequence work. A model that produces gorgeous standalone clips can be the wrong pick for a dialogue scene that must stay consistent for thirty seconds. Decide which bucket you are in before you open a single prompt field.

Choosing a Model by Shot Type, Not by Hype

Model coverage changes fast, and brand-name ranking is a poor decision rule. A more durable approach: classify the models by the behavior they are optimized for, then map your shot to the behavior.

The cinematic generalists. This category — the big flagship generators — excels at photoreal texture, believable lighting, and complex camera language. Lens flares behave. Depth of field falls off correctly. These are the right tools for hero shots, product beauty shots, and any frame that will be paused and inspected. Their weakness is cost per usable second and, historically, weaker fine control over exact composition.

The efficiency specialists. A second family of models, including several strong Chinese and Korean-built systems, optimized for different trade-offs: fast iteration, lower compute per second, and tight stylistic range. They can be superb for stylized animation, social-first vertical formats, and high-volume variant testing where you need twenty options, not one perfect option. Loop through several of these when you are exploring a look, then move the winning direction to a cinematic generalist for the final render.

The motion-first models. A third family is built around movement rather than scenery: the visual-motion models. These are the right call when the shot is defined by an action — a whip pan, a character turning, fabric in wind, a hand opening a door. If you describe motion and the output arrives static, you used the wrong family.

The image-animation tier. Some systems are designed specifically to take a still and bring it to life with controllable directional motion. If your workflow starts from stills — storyboards, product renders, illustrations — build your pipeline around these rather than forcing a text-only model to hallucinate your existing asset.

Practical decision rule: write down the shot in one line, underline the noun that matters, and underline the verb that matters. Noun-dominant shots want a cinematic generalist. Verb-dominant shots want a motion-first model. Volume work wants an efficiency specialist.

Writing Prompts That Behave Like Shot Lists

The single biggest quality jump comes from restructuring prompts. Freeform adjective soup produces mush. A shot list produces footage.

A workable prompt skeleton, in this order:

  1. Shot size and angle — wide establishing, medium two-shot, extreme close-up, low-angle hero shot.
  2. Subject and action — one subject, one primary action, stated in present tense.
  3. Setting and time of day — location, weather, light direction.
  4. Camera behavior — static locked-off, slow dolly in, handheld follow, crane up.
  5. Lens and rendering — 35mm shallow depth of field, anamorphic flare, 24fps motion cadence.
  6. Look and grade — cool teal shadows, warm practical highlights, high-contrast noir.
  7. Continuity anchors — costume, hair, prop, and any element that must match the neighboring shot.

Compare two prompts for the same idea.

Weak: "A sad man in a city at night, cinematic, amazing, highly detailed, beautiful lighting, masterpiece."

Structured: "Medium close-up, low angle. A man in his forties in a wet wool coat stands at a crosswalk, rain falling through a passing car's headlights. He exhales and looks down at his hands. Camera handheld, slow push in. 50mm, shallow depth of field, cool shadows with warm sodium highlights. Must match: grey coat, silver-rimmed glasses, no umbrella."

The second prompt gives the model triage instructions. When compute is limited, it knows what to protect: the face, the coat, the light direction. The first prompt gives it nothing to prioritize, so it invents.

Two disciplines matter more than any single phrasing trick. First, one action per clip. Models degrade when asked to perform a sequence of verbs in a few seconds. Second, state what stays still, not just what moves. "Static camera" and "background unchanged" are legitimate, useful instructions.

A Five-Step Production Loop You Can Repeat

Generating clips randomly is how people burn a week. A repeatable loop looks like this.

Step 1 — Beat the script down. Take your scene and reduce it to one line per shot: what changes on screen in each. If two shots change the same thing, merge them.

Step 2 — Lock the look on stills. Generate or select key frames before you generate anything moving. Stills are cheap to iterate and reveal composition problems instantly. Choose one still per shot as your anchor.

Step 3 — Animate from the anchor. Feed the anchor into an image-to-video or motion-driven pipeline, with a prompt describing only motion and camera, not appearance. Appearance is already in the frame. Restating it wastes prompt budget and invites drift.

Step 4 — Generate variants, not retries. When a take fails, change one variable, not five. Same prompt, new seed. Then same seed, shortened action. Then same seed, stronger camera instruction. If you change seed, prompt, and model simultaneously, you learn nothing.

Step 5 — Assemble and gate. Cut the takes together before polishing. A clip that looks stunning alone can fail in a sequence due to mismatched motion cadence or light direction. Fix sequence problems now, not in the final grade.

This loop exists because generation is stochastic. You are sampling from a distribution. Good directors control the sampling instead of rerolling blindly.

Continuity: The Hardest Problem in AI Video

Individual clips have gotten good. Sequences have not gotten easy. Here is where to spend your attention.

Character consistency. Anchor on a single reference image and reuse it. Prefer image-conditioned generation over re-describing a face in text, because language underspecifies faces. Keep a written continuity sheet for each character: hair, wardrobe, distinguishing marks, and any prop they carry. Paste that sheet into every prompt for that character, verbatim. Consistency comes from repetition of the same words, not from new descriptive flourishes.

Environment consistency. Reuse the establishing shot's anchor frame as a reference for the reverse angle. Keep light direction fixed across the scene — if the sun is camera-left in the wide, it stays camera-left in the close-up. Most "AI looks wrong" complaints in sequences are actually lighting continuity errors.

Motion cadence. Two clips at the same nominal frame rate can feel different if camera movement speeds differ wildly. Decide on a movement grammar for your scene: is the camera drifting, or locked? Enforcing one rule across a scene hides a lot of small imperfections.

Multi-asset quality control. When a project has many generated assets — backgrounds, characters, props, transitions — quality drifts silently. Run a checklist pass: resolution match, color temperature match, grain match, frame rate match, and a final viewing at 100% zoom for artifacts. Cheaper to catch ten mismatches in a contact sheet than one in a client review.

A useful technique borrowed from pixel-art and mosaic workflows: assemble a contact sheet or coarse composite of every asset in the sequence at low resolution before rendering anything at full quality. At small size, structural problems — a character whose shoulders sit too low, a prop that changed color, a background that reads too dark — become obvious. This mirrors the discipline behind grid-based art techniques where every tile must sit correctly within a larger pattern. Iterate on the composite until the whole thing reads well, then commit to high-resolution passes.

Directing the Edit Before the Edit

An underrated advantage of generative footage: you can plan the cut in the prompt.

Shoot coverage. Generate more than you need from multiple angles. An over-the-shoulder, a wide, and an insert of the hands costs little and saves a scene that does not cut. Editors have been doing this for a century; generative work should not abandon it.

Generate matching eye-lines. If shot A has a character looking frame-right, shot B's reverse should have them looking frame-left. Specify gaze direction explicitly. It is one of the most common continuity mistakes in generated sequences and one of the easiest to prevent.

Build transition shots on purpose. A close-up of a hand on a door handle, a shot of a flickering sign, a shot of feet on gravel — these are edit glue. Generate glue shots in batches and keep them in a library. When a cut feels abrupt, a two-second insert usually fixes it.

Match action across cuts. If a character raises a glass in shot A, shot B should begin with the glass already raised. Prompt the end state of the previous shot as the start state of the next.

Where Quality Usually Falls Apart

A debugging guide, in rough order of frequency.

Melted hands and faces. Cause: too much action in too few seconds, or a subject occupying too little of the frame. Fix: shorten the action, move the camera closer, or split into two shots.

Unwanted style shifts mid-clip. Cause: conflicting style adjectives. Fix: pick one grade and one lens language per clip, and remove decorative words.

Character changes clothes or hair. Cause: appearance described differently between prompts. Fix: verbatim continuity sheet, pasted unchanged every time.

Motion that reads as video-game-like. Cause: camera instruction too vague or physically impossible. Fix: name a concrete camera rig behavior — dolly, crane, handheld, locked-off — and remove multiple simultaneous moves.

Muddy, oversharpened output. Cause: over-stacking quality buzzwords like "8K, ultra detailed, masterpiece, hyperrealistic." These often fight each other. Fix: describe the scene, then add at most one or two technical rendering notes.

Good clips, bad scene. Cause: no coverage, no continuity plan. Fix: anchor frames, contact sheet review, and a written shot list.

Building a Small, Durable Pipeline

Tools churn. A pipeline outlives them.

Keep four libraries. A prompt library of tested shot templates by shot type. An anchor library of approved stills per character and location. A glue library of short inserts and transitions. And a rejection log recording what failed and why — this is the asset nobody builds and everybody needs, because it is the only thing that stops you repeating the same mistake across projects.

Standardize your outputs for the edit. Pick one resolution and frame rate for a project and convert anything that deviates on import. Keep generated clips at the highest quality your storage allows and do the grading once, at the end, on the assembled timeline.

Finally, budget your generation time like a shoot day. Reserve roughly a quarter of your time for exploration, half for the shots that need to be right, and a quarter for re-shoots caused by continuity failures. Teams that plan for re-shoots ship scenes. Teams that assume first takes will all land do not.

The Skill That Transfers

Every new model release makes this question louder: what is the human still for? The answer has not changed much in a hundred years of filmmaking. Someone has to decide what the scene is about, what the audience should feel, and what must not change from one moment to the next. Models execute. Directors specify.

That is the real reason shot lists, continuity sheets, coverage, and anchor frames matter more than any single prompt phrase. They are the vocabulary that converts intent into footage. A creator who can hand a system a precise, prioritized, physically plausible shot description will outperform someone with access to better models and no plan — reliably, and increasingly so as generation quality converges.

Text-to-video is not a creativity machine. It is a camera that needs a director. Learn to specify well, and the tool stops feeling like a slot machine.

FAQ

How long should a generated clip be?
Short. Four to eight seconds is the sweet spot for most models, because longer durations invite drift and invention. Build sequences from short shots, the way real scenes are cut.

Should I write prompts in my native language or in English?
Follow the model's training strength and test both for a specific shot. If your output needs to be evaluated in English or the model's documentation is English-first, English prompts often behave more predictably. What matters more is using one consistent vocabulary.

Do I need a reference image for every shot?
No, but you need one for anything that must recur. Establish a character or location with a reference, then reuse it. For one-off shots, a strong written prompt is enough.

How many variants should I generate per shot?
Two to four well-controlled variants beat twenty random rerolls. Vary one variable at a time so you can attribute what improved.

How do I stop characters from changing between shots?
Use a fixed reference image, keep a verbatim continuity description for each character, fix light direction across the scene, and review the whole sequence at low resolution before committing to high-quality renders.

Is it better to generate stills first, or go straight to video?
Stills first. They are faster and cheaper to iterate, they reveal composition problems immediately, and they double as anchors for image-conditioned generation.

What causes that plastic, artificial look?
Usually over-stacked quality keywords, a vague camera instruction, and too much simultaneous action. Simplify the prompt, name a real camera behavior, and shorten the action.

How do I plan a scene instead of a clip?
Write one line per shot, define what changes in each, generate coverage from multiple angles, match eye-lines and action across cuts, and keep a library of short insert shots for transitions.

Alexander

Alexander