Why AI Video Needs Direction, Not Just Prompts
A single generated shot can look astonishing. Ask a model for a rain-soaked street at night and you will get atmosphere, reflections, and motion that would normally take a crew hours to light and shoot. Ask for that same street across four connected shots and the seams appear immediately: the rain changes density, the storefront signs mutate, the actor's coat shifts from charcoal to navy, and whatever story you imagined dissolves into a very pretty mood board.
The difference between a clip and a sequence is direction. A director decides what the audience should feel at each moment, what information they need in order to keep following the story, and which details must remain stable so the story reads cleanly. In a traditional production, those decisions live in the script supervisor's notes, the continuity photographs, the lighting plan, and the edit bay. In an AI pipeline, they live in the documents you write before you open a generator and in the discipline you apply while prompting.
This guide describes a complete script-to-screen workflow for AI video. It is deliberately engine-agnostic. You can run it with browser-based image-to-video tools, a local diffusion setup, or a hybrid pipeline that mixes several engines and a traditional editor. Tool names change every few months. The discipline does not.
One more framing point before the mechanics. Most people who struggle with AI video are not struggling with technology. They are struggling with authorship. They have a vague feeling about a mood, a playlist, and a reference image, and they expect a text box to resolve the ambiguity. Generators amplify ambiguity. They never resolve it. Everything below is a method for removing ambiguity before it becomes expensive.
The Five Documents Every Project Needs
You do not need a production binder. You need five short documents, each of which answers a question the next stage will ask. Write them in order and resist the urge to skip ahead to prompting.
The single-sentence spine
Write one sentence that states what changes between the first frame and the last. "A cautious hiker learns that the shortcut is worth the risk." "A night-shift baker realizes the regular who never speaks has been leaving notes in the bread bags." If you cannot produce that sentence, no model will rescue the project, because the model has no idea what it is building.
The spine acts as a decision filter for the rest of production. When you are choosing between a wide shot and a close-up, you are really asking which one advances the change described in that sentence.
The beat sheet
Break the spine into four to eight beats, each with an approximate duration. A 45-second film usually holds five to seven beats. Beats are emotional or informational units, not shots. "She hesitates at the fork" is a beat. "Medium shot of boots on gravel" is a shot. Keeping those two categories separate is what prevents you from generating footage with no dramatic function.
The screenplay with visual intent
Write the dialogue, narration, or on-screen text first, then add visual notes in plain language. This is the document you will actually read while prompting, so it should describe intent rather than camera jargon you cannot reliably reproduce. "We need to feel the wind before we see the person" is more useful than "whip pan to reveal."
The shot list
Convert beats into shots. For each shot, record a number, an estimated duration, framing, subject action, environment, lighting direction, and the emotional job the shot performs. That final column is what separates a director's list from a shopping list.
The asset and continuity sheet
A running list of every recurring element: character reference images, wardrobe descriptions, location phrases, color palette, lens character, and reusable seeds or settings. Anything you intend to repeat across shots belongs here, spelled exactly the same way every time.
Writing a Screenplay a Video Model Can Execute
Screenplays written for humans and screenplays written for generators are not identical documents. Both need clarity, but generators reward a specific kind of clarity: concrete nouns, physical actions, and a single dominant idea per moment.
Write dialogue and narration first
Voice performance drives pacing more than any visual decision. If narration runs long, every shot after it gets stretched or rushed to fit, and the edit becomes a negotiation with audio instead of a creative choice. Write the lines, read them aloud with a stopwatch, and cut until the rhythm feels natural at roughly 140 to 150 words per minute. Synthetic voices rush when you deny them punctuation, so give them commas, short sentences, and full stops to breathe on.
Add visual intent, not camera jargon
For each line of dialogue, note what the audience should be looking at and why. "While she says she is fine, we watch her hands" is direction. "Dutch angle, 18mm, rack focus" is a technical wish that most generators will partially ignore. You can add technical notes later, once you know which engine you are using and what it can actually control.
Scene economy: 45 seconds is a whole film
A 45-second piece should contain roughly the amount of story you would put in a single scene of a longer work: one location, one or two characters, one change. If your outline needs three locations and a time jump, either extend the runtime or accept that the result will feel like a trailer. Trailers work, but only when they are built on purpose with title cards, voiceover, and a strong music bed.
Read it aloud before you generate anything
Record yourself reading the screenplay with a phone. If your delivery sounds like a person telling a story, the script is ready. If it sounds like a list of shots, it is not a script yet. This five-minute test saves hours of generation on material that was never going to cut together.
From Beats to Shot List: Building the Directing Grid
A shot list is a table, and a good table is boring in the best way. Here is a compact example for a 20-second excerpt.
| Shot | Duration | Framing | Action | Environment | Light | Emotional job |
|---|---|---|---|---|---|---|
| 1 | 4s | Wide | Hiker walks along ridge line | Granite ridge, cloud valley below | Cold blue dawn from camera left | Establish scale and solitude |
| 2 | 3s | Medium tracking | Boots and lower legs on gravel | Same trail surface | Same direction, softer contrast | Forward motion, effort |
| 3 | 2.5s | Close-up | Hand pulls a compact flask from a pack | Canvas pack, worn straps | Same direction, shallow depth | Introduce the object |
| 4 | 3s | Insert | Steam rises against cold air | Flask mouth, dark background | Backlit rim light | Sensory payoff |
| 5 | 4s | Wide | Hiker sits facing the valley | Same ridge, slightly different angle | Same direction, wide tonal range | Stillness |
The grid answers questions that would otherwise be answered by accident during generation. How long is the shot? What is in frame? Which way is the light coming from? What is the shot for?
The emotional job column is not optional
When a shot's purpose is unclear, you will over-generate. You will produce nine variations of a handsome image and then discover in the edit that none of them advance the story. Naming the job — establish, orient, reveal, react, resolve — tells you when a shot is finished.
Do the coverage math early
A reliable planning ratio is roughly one shot per three seconds of finished runtime, plus about 30 percent extra coverage for safety. A 45-second film therefore wants around 15 planned shots and about five spare options. Coverage is what lets you fix pacing during the edit instead of going back to the generator with a stopwatch and a deadline.
Prompt Structure: Six Slots, Style Locks, and Negative Direction
Once the shot list exists, prompting becomes transcription rather than improvisation. That is exactly the goal: you want the creative decisions to have been made in a calmer state than the one you will be in at two in the morning with a render queue open.
The six-slot frame
Write every prompt in the same order, covering six slots: subject, action, environment, lighting, camera, and mood or style. Keeping the order fixed makes it easy to spot what changed between two attempts.
- Subject: who or what, including wardrobe and physical detail. "A weathered climber in a red shell jacket" beats "a man."
- Action: one clear physical verb. "Tightening a boot strap" beats "getting ready."
- Environment: location, weather, surface, background texture.
- Lighting: direction, quality, time of day, color temperature.
- Camera: distance, movement, lens feel, framing logic.
- Mood and style: tonal reference, palette, grain, finish.
A finished example: "A weathered climber in a red shell jacket, tightening a boot strap, on a granite ledge above a cloud-filled valley, cold blue morning light from camera left, slow handheld medium close-up, restrained documentary tone."
Style locks keep a sequence looking like one film
A style lock is a short phrase you repeat verbatim in every prompt of a sequence: "shot on 35mm, shallow depth of field, muted teal and rust palette, fine grain." If the phrase drifts — "muted teal" in one shot and "moody blue" in the next — the grade will drift too, and you will spend the edit trying to unify footage that should have been consistent from the start. Paste the lock rather than retyping it; typos are a real source of drift.
Negative direction is part of the prompt
Most engines accept exclusions, and a short reusable list pays for itself: no text overlays, no watermarks, no extra limbs, no fast whip pans, no lens flare, no cartoon shading, no duplicated faces. Keep the list short. A bloated negative list starts contradicting itself and can suppress the very qualities you asked for in the positive prompt.
Specificity beats poetry, and brevity beats both
Words like "epic," "cinematic," and "breathtaking" carry almost no instruction. "Low angle, 24mm, subject centered, horizon one third from the top" gives a model something it can obey. Keep prompts under roughly 80 words. Long prompts dilute the concepts that matter most, because every added clause competes for the same limited attention.
The three most common prompt failures
The first is changing two variables at once. If you alter the lighting and the framing together and the shot improves, you have learned nothing reusable. Change one slot per attempt.
The second is reusing a prompt across shots that need different intent. Copy-paste is efficient for style locks and wardrobe, and disastrous for action and framing.
The third is describing the story instead of the shot. "She realizes she has been wrong" is not visible. "A slight pause, eyes lifting from the map to the horizon" is.
Character and Location Consistency Without Reshoots
Consistency is the most common quality complaint in AI video, and it is solved mostly by process rather than by finding a better engine. Treat it as a continuity department that you staff yourself.
- Build a character sheet: one front-facing portrait, one three-quarter view, one profile, one full body. Use those images as references in every shot where the character appears.
- Freeze wardrobe in writing, then never change a word of it. "Olive canvas jacket, brass zipper, no hat" stays identical across all prompts.
- Lock locations with a single approved establishing image, and repeat the same environment phrase in every prompt, including the words you would normally vary for elegance.
- Record seeds and settings beside the shot number when the engine supports them. Reproducibility is continuity insurance.
- Match lighting direction across a scene. A character lit from the left in shot three and from the right in shot four reads as a different time of day, even if everything else matches.
- Keep color temperature consistent. Mixing a warm interior with a cool exterior inside one scene is a deliberate choice, not an accident to discover in the edit.
When a shot refuses to behave after three attempts, change the shot rather than fight the model. Cut to hands. Turn the camera away. Frame the character from behind. Directors solve continuity problems with coverage, not stubbornness, and a clever insert is almost always cheaper than a perfect face.
Choosing an Engine Shot by Shot
Not every shot deserves the same tool, and the best pipeline is usually a mix. Judge each shot against a short list of criteria.
| Criterion | Question to ask | Practical implication |
|---|---|---|
| Motion complexity | Does the camera move, the subject move, or both? | Complex camera paths favor engines with explicit camera control |
| Identity preservation | Does a recurring character appear? | Requires image reference support and a locked wardrobe phrase |
| Duration | Is the shot longer than five seconds? | Plan to extend, stitch, or split into two shots |
| Audio needs | Dialogue, ambience, sync? | A separate audio pipeline is usually more controllable |
| Stylization | Photoreal or illustrative? | Some engines excel at one and merely tolerate the other |
| Iteration speed | How many attempts can you afford? | Fast drafts beat slow perfection on uncertain shots |
| Output format | Vertical, square, or widescreen? | Compose for the delivery ratio from the first render |
The practical rule is simple: explore with your fastest option and finish with your most controllable one. Never explore with your slowest, most constrained tool, because exploration requires volume. And never finish a hero shot with a tool that cannot hold an identity for four seconds.
When to stitch instead of extend
If a shot needs eight seconds of screen time but your engine holds coherence for four, do not fight it. Design the shot as two connected framings: a medium that pushes toward the subject, then a close-up that completes the action. Cutting between them hides the seam and often improves the pacing anyway. Audiences accept a cut far more readily than they accept morphing flesh.
Keep a small render log
For each shot, note the engine, the prompt version, the seed, the number of attempts, and a one-word verdict. Two projects in, this log becomes your most valuable asset, because it tells you which tool actually solves which problem in your specific style.
Pacing, Coverage, Editing Rhythm, and Sound
AI clips frequently look better than they cut, because each clip is a complete, self-satisfied moment. Real sequences breathe unevenly. A few rules keep a generated film from feeling like a slideshow of beautiful mistakes.
Shot length and cut points
An average shot length between 2.5 and 4 seconds maintains energy without feeling frantic. Cut on motion: trim a couple of frames into a movement and the seam disappears into the gesture. Give the audience one idea per shot; two subjects doing two things splits attention and reads as noise. Insert a wide shot every four or five tight shots to reset spatial understanding, and hold the final frame of a scene long enough for the viewer to register that something changed.
If a cut feels slow, shorten shots rather than speeding them up. If it feels incoherent, you are probably missing an establishing shot or a reaction shot, not a transition effect.
Build the sound track in layers
Viewers forgive imperfect motion far more easily than bad audio. Layer the track in this order: voice, then room tone or ambience, then music, then spot effects. Generating or recording voice first is important because pacing decisions in the edit depend on it — you cannot time a cut to a line you have not heard.
Music should follow the beat sheet rather than the runtime. Let the theme enter on the decision beat and resolve on the final beat. For a 45-second piece, a single sustained drone plus one percussive hit often outperforms a full orchestral bed, because it leaves space for the narration and the ambience.
Do not let the generator score your film
If an engine produces audio alongside video, treat that audio as scratch material. Replace it, or at least audition the result with headphones and a critical ear for artifacts. The most common giveaway of a generated film is not the visuals; it is thin, phasey, oddly gated sound.
Worked Example: A 45-Second Ridge Coffee Film
Spine: "A solo hiker discovers that the best part of the trail is the pause." Runtime target: 45 seconds. Deliverable: vertical 9:16 for social, with a widescreen master for the portfolio.
- Wide establishing shot, dawn ridge, 4 seconds — scale and solitude.
- Medium tracking shot, boots on gravel, 3 seconds — forward motion and effort.
- Close-up, hand pulls a compact flask from a pack, 2.5 seconds — the object enters the story.
- Insert, steam rising against cold air, 3 seconds — sensory payoff.
- Wide, hiker seated facing the valley, 4 seconds — stillness and reward.
- Close-up, slight smile and exhale, 3 seconds — emotional resolution.
- Slow pull-back wide, 5 seconds, with title card — thesis: the pause is the point.
That is roughly 24 seconds of planned footage expanded to 45 seconds with pacing, sound design, a title card, and held frames. This ratio is normal. Do not expect generated footage to fill the runtime one to one.
Production notes from the run: all seven prompts reused the same six-slot frame, the same "cold blue morning light from camera left," and the same jacket and flask descriptions. Two character reference images covered shots two, five, and six. Shot four needed three attempts before the steam read as steam rather than smoke, and the winning version used a backlit rim light clause that the earlier attempts lacked. The edit trimmed each shot on motion, brought the music in at shot three, and held two extra frames on the final wide before the title faded up.
The lesson worth keeping: the shots that fought back were the ones whose intent was vague in the list. Once shot four was rewritten as "we need to see that the air is cold," the prompt practically wrote itself.
Quality Control, Common Mistakes, and FAQ
The three-pass review
Watch the finished piece three times with three different kinds of attention. First pass: does the story land? Ignore everything else and ask whether a stranger would understand the change from first frame to last. Second pass: look only at hands, faces, and any on-screen text, which is where generation artifacts hide. Third pass: listen with your eyes closed and note where the audio loses energy or where a line is hard to parse.
Fix the top three problems and stop. Perfectionism is the enemy of shipping, and the fourth problem you would have fixed is almost never visible to anyone but you.
The pre-publish checklist
- Loudness is normalized to a consistent level across the whole piece.
- Captions or subtitles are present, correctly timed, and inside safe margins.
- The thumbnail frame is deliberate, not an arbitrary scrub position.
- The first two seconds contain motion, a face, or a line of text that earns attention.
- Every asset, voice, and music track is cleared for the way you intend to publish.
- A high-bitrate master is archived separately from the platform export.
Mistakes that sink generated films
- Writing the video in a prompt box instead of on paper. If the story exists only in a text field, it will not survive the edit.
- Chasing a perfect shot indefinitely. Three failed attempts usually mean the shot is wrong, not the model.
- Changing style words between shots. Consistency is a vocabulary habit, not a rendering feature.
- Ignoring the delivery aspect ratio until the end. Cropping a composed vertical shot to widescreen destroys the framing you paid for.
- Treating audio as an afterthought. Weak sound makes excellent footage feel amateur.
- Skipping the rough cut. You cannot judge generated footage until it is cut against other footage.
- Overloading motion. Slow, deliberate movement reads as professional; constant movement reads as synthetic.
- Generating without a shot number in the filename. Two days later, you will not know which file belongs to which beat.
- Forgetting that a scene needs a reaction shot. Audiences follow emotion more reliably than they follow action.
- Publishing the first export. Watch it once on a phone, once on a laptop, and once with headphones before you upload.
FAQ
Do I need a full screenplay for a 15-second clip?
You need a spine and two beats. Even a very short piece reads better when you know what changes between the first and last frame, and a spine takes ninety seconds to write.
How many shots should I generate per finished second?
Plan roughly one shot per three seconds of runtime, plus about 30 percent extra coverage. The extra material is what lets you solve pacing problems in the edit instead of returning to the generator.
Can one engine handle the entire project?
Sometimes, but mixing per shot is normal and healthy. Use the most controllable option for hero shots and the fastest option for connective shots, and keep a render log so you remember which is which.
How do I stop characters from changing between shots?
Reference images, frozen wardrobe language, consistent lighting direction, consistent color temperature, and reusable seeds. If a shot still drifts after three attempts, change the coverage rather than the model.
Is it better to generate video from text or from an approved still?
Approved still first, almost always. Approving a keyframe catches composition, identity, and wardrobe problems before you spend time and processing on motion.
What resolution and frame rate should I export?
Match the destination platform, deliver 24 or 30 frames per second for a cinematic feel, and always archive a high-bitrate master for future edits and re-cuts.
How long does a 45-second piece take?
With a locked script and shot list, expect a few hours of generation, an hour of assembly, and an hour of sound work. Most of that time goes into iteration, which is precisely what the shot list exists to contain.
What if my footage looks good but the film still feels flat?
Flatness is almost always a story problem rather than an image problem. Return to the spine and ask what genuinely changes. If nothing does, the piece is a montage, and montages need music and pacing to carry them rather than character.
Start with one scene, not one film. Pick a single beat, run the full pipeline end to end, and keep the shot list, the prompts, and the render log from that run. The second scene will take half the time. By the tenth, prompting stops feeling like gambling and starts feeling like directing — which is the whole point of writing the documents first.



