Why the Prompt Is the New Screenplay
A short film used to begin with a screenplay: slug line, action, dialogue, and a director's interpretation layered on top. In an AI-assisted pipeline, the same idea begins with a prompt, but that prompt has to carry far more weight. It describes the subject, the lens, the light, the movement, the pacing, and the emotional temperature all at once, in a form a generative model can act on. Get it right and a single paragraph becomes six seconds of convincing cinema. Get it vague and you spend an afternoon rerolling clips that look almost right.
The craft is not about memorizing magic words. It is about learning to think like a director and write like a technical editor at the same time. A prompt is a specification: it defines what must be in frame, what must not be, and how the frame should behave over time. The more precisely you can describe an image you already see in your head, the less the model has to guess on your behalf.
There is also a psychological shift worth naming. When you write a screenplay, you hand off interpretation to a crew. When you write a prompt, you are the crew. You are the cinematographer choosing the lens, the gaffer choosing the key light, the editor deciding where the cut lands. That is more responsibility, but it is also more control than most first-time filmmakers have ever had.
This guide walks through a complete pipeline for turning a rough idea into a finished short film: structuring prompts, planning shots, maintaining continuity, matching prompts to different classes of model, and assembling everything into something that holds together for a minute or two. It assumes no budget, no crew, and no specialized hardware, just a clear idea and a willingness to iterate.
The Anatomy of a Cinematic Video Prompt
Most weak prompts fail for the same reason: they describe a subject and stop. "A woman walking through a rainy city at night" is a subject, not a shot. A strong prompt stacks six layers, and each layer answers a question the model would otherwise answer randomly.
Layer 1: Subject and Action
Be specific about age range, wardrobe, posture, and emotional state. "A woman" becomes "a woman in her late thirties wearing a damp olive raincoat, shoulders hunched, jaw tight." Give her one primary action with direction: she exhales, she turns left, she lifts a paper cup. One action per shot is a useful discipline. Two actions usually produce a mushy compromise where neither reads clearly.
Layer 2: Camera and Framing
Shot size (extreme wide, wide, medium, close-up, macro), angle (low, eye level, high, dutch), and lens feel (24mm for environmental context, 50mm for natural perspective, 85mm for compression) do enormous work. Add the camera support too: handheld for tension, gimbal for elegance, locked-off tripod for stillness, drone for scale. Composition hints such as centered symmetry or negative space to the left help the model place the subject rather than centering everything by default.
Layer 3: Light and Color
Describe the direction, quality, and source of light: soft window light from the left, hard sodium-vapor practical from behind, overcast diffusion with no visible shadow. Then commit to a palette. "Cool blue shadows with amber highlights" is a decision. A vague "nice lighting" is not.
Layer 4: Motion and Timing
Separate subject motion from camera motion and state both. A slow push-in, a gentle handheld drift, rain falling, steam rising. Specify whether the shot is real time, slow motion, or a time-lapse, and roughly where it begins and ends. This layer is where most jitter and morphing artifacts are prevented, because the model understands it is meant to move smoothly rather than invent transitions.
Layer 5: Style and Medium
Photorealistic, documentary handheld, 16mm grain, archival footage, stop-motion, cel animation, 3D render. Pick your style words once and repeat them in every prompt of the film. Style consistency is the single cheapest way to make separately generated clips feel like one project rather than a reel of unrelated experiments.
Layer 6: Constraints and Negative Instructions
Say what you do not want, but keep it short and concrete: no text overlays, no warped hands, no watermark, no sudden cuts, no camera shake. Long negative lists tend to fight the rest of the prompt and can flatten the image. Three to six constraints is usually plenty.
Here is the whole stack assembled into one prompt:
Medium close-up, handheld, a woman in her late thirties in a damp olive
raincoat stands at a bus stop, rain streaking the glass behind her, she
exhales slowly and looks left off-frame, sodium-vapor practical light from
the right, cool blue shadows, shallow depth of field, 85mm compression,
16mm grain, slow breathing motion only, no text, no watermark.
Aim for roughly 40 to 90 words per shot. Front-load the most important elements, because attention tends to fade toward the end of a long prompt.
From Logline to Shot List in Four Passes
Planning is what separates a film from a folder of clips. Four passes is usually enough.
Pass one: the logline. One sentence with a character, a want, and an obstacle. "A night-shift nurse misses the last bus and has to cross a flooded city on foot before dawn." That sentence is your north star. Every prompt decision should serve it.
Pass two: the beat sheet. Five to eight beats, each one a change. She finishes her shift. She finds the bus gone. She starts walking. She meets a stranger. She arrives too late. Beats give you the shape; shots give you the texture.
Pass three: the shot list. For a 60- to 90-second film, plan 8 to 14 shots. Write each as a single line: wide establishing shot of empty bus stop in rain; close-up of her hand checking a dead phone; medium tracking shot behind her as she walks through a flooded underpass. Note the duration you expect each to hold on screen, typically two to four seconds.
Pass four: prompt cards. Turn each shot-list line into a fully layered prompt using the six layers above. Keep them in a single document in shot order. This document is your film. Everything else is execution.
A shot list also protects you from a common trap: generating gorgeous clips with no narrative through-line. If a shot does not advance a beat, cut it before you spend time generating it.
Engineering Continuity Across Shots
Audiences forgive imperfect images but not broken continuity. If your lead's coat changes color between shots, the illusion collapses immediately. Continuity needs to be engineered deliberately.
Lock a character block. Write a short, fixed description of each recurring character and paste it verbatim into every prompt that features them. Same age, same hair, same wardrobe, same adjectives in the same order. Small wording changes produce visible changes in output.
Use reference stills. Generate a strong still of your character first, then drive subsequent shots with image-to-video so appearance is anchored by the image rather than by text alone. This is the most reliable method for keeping a face consistent.
Anchor to props and wardrobe. Distinctive elements, a red scarf, a cracked phone, a canvas bag, give the viewer something to track and give the model a memorable token to reproduce.
Maintain screen direction. If your character exits frame right, she should enter frame left in the next shot. Reversing this reads as a jump in space and disorients the viewer even if they cannot say why.
Keep light and time of day stable within a scene. Write the scene's lighting conditions once and reuse them. Scenes that drift from golden hour to noon and back destroy the sense of a single continuous moment.
A simple continuity sheet with columns for character, wardrobe, location, time of day, palette, and lens style takes ten minutes to build and saves hours of regeneration.
Matching Prompts to Different Video Models
Not every engine wants the same prompt shape. Writing for the wrong class of model is a frequent source of frustration.
Text-to-video models reward complete, descriptive prompts that include appearance, action, and camera. They are excellent for establishing shots, landscapes, and abstract imagery, but weaker at precise identity continuity across shots.
Image-to-video models treat the input frame as the answer to "what does this look like." Your prompt should therefore focus almost entirely on motion, camera, and timing, and should avoid re-describing appearance, which can cause the model to fight its own reference image.
Motion-focused tools often respond best to short prompts built around a clear movement verb and a direction: slow push in, gentle orbit right, rack focus from foreground to background. Adding a paragraph of style text here can dilute the motion instruction.
Dialogue and lip-sync tools need the line text, the speaking pace, head movement instructions, and an almost static camera. Wide gestures and moving cameras are the enemy of believable sync.
Decision criteria are simple. If the shot's value is in its precise look, start from a still and drive motion with image-to-video. If the shot's value is in complex movement through a space, start with text-to-video and accept looser identity control. If the shot needs a face to carry emotion, generate fewer, longer, better-anchored takes rather than many short ones.
A Repeatable Production Workflow
Here is a workflow that scales from a 30-second test to a five-minute short.
Step 1: Write the One-Page Treatment
Half a page of prose describing the story, the mood, and the visual rules. No prompts yet. This keeps the project anchored to intent rather than to whatever looks cool.
Step 2: Build a Look Bible
Decide aspect ratio, palette, lens character, grain, era, and reference films. Write these as reusable phrases. Every prompt you write afterward includes them.
Step 3: Board Shots as Stills
Generate still images for each shot before generating any video. Stills are faster and cheaper to iterate, and a shot that does not work as a still will rarely work as motion.
Step 4: Write Prompt Cards
One card per shot with all six layers, the character block, the look bible phrases, and the intended duration. Keep cards in order in a single document.
Step 5: Generate in Batches
Run three to five variations per card. Change one variable at a time when iterating, so you learn what actually caused the improvement.
Step 6: Apply the Three-Take Rule
If a shot has not worked after three attempts, the prompt is wrong, not your luck. Rewrite the card, simplify it, or change the approach entirely, for example switching from text-to-video to a still-driven shot.
Step 7: Cut to a Scratch Track
Assemble the best clips against a temporary music bed or a click track before you polish anything. Editing exposes missing coverage immediately, and it is much cheaper to discover that now than after upscaling.
Step 8: Finish
Upscale, interpolate frames if motion looks choppy, color match clips to a common grade, add ambience and foley, and export. Save each stage separately so you can return to a clean intermediate if something breaks.
Adopt a naming convention such as shot03_v2_select and keep a text log of which prompt produced which take. Without versioning, you will rediscover a good prompt by accident and never be able to repeat it.
Common Prompting Mistakes and Fixes
| Mistake | What it looks like | Fix |
|---|---|---|
| Writing a story instead of a shot | Prompt covers three actions and a location change | One action, one shot, one camera move |
| Contradictory style words | "Photorealistic anime watercolor" | Choose one medium and repeat it |
| Overloaded negatives | Ten "no" clauses that flatten the frame | Keep three to six specific constraints |
| No motion instruction | Subject drifts or morphs unpredictably | State subject motion and camera motion separately |
| Crowded frames | Multiple characters whose faces all degrade | One or two subjects, wider shots for crowds |
| Inconsistent descriptions | Character changes between shots | Use a fixed character block and reference stills |
| Chasing perfection in generation | Dozens of takes for one shot | Fix in the edit; cut around weak frames |
| Ignoring duration and aspect ratio | Clips that cannot be cut together | Specify both in every prompt |
| Vague emotional adjectives | "Epic and beautiful" with no visual meaning | Translate emotion into light, color, and framing |
The last row deserves emphasis. Words like epic, dreamy, and cinematic are not instructions; they are feelings. Convert them. Cinematic might mean anamorphic flare, shallow depth of field, and a slow dolly. Dreamy might mean diffused highlights, slight overexposure, and drifting handheld motion. Once translated, the model can act.
Sound, Pacing, and the Final Assembly
Generated visuals carry a slight uncanniness that sound design erases. A room tone bed, a few well-placed foley hits, and a music cue that enters on the right beat make an audience accept imagery they would otherwise question. Budget as much time for audio as for generation.
Pacing is where AI shorts usually fail. Beginners hold shots far too long because each clip took effort to produce. Cut on movement, keep most shots between two and four seconds, and vary the length deliberately: a run of short shots builds tension, one long shot releases it. If a clip only looks convincing for two seconds, show it for two seconds.
Color is the other great unifier. Generated clips arrive with slightly different contrast, saturation, and white balance. Apply a common grade across the whole film, match your blacks and highlights, and consider a subtle grain layer over everything. A uniform look does more for perceived production value than any individual shot.
Finally, structure the ending. AI short films tend to stop rather than conclude. Decide on a final image that answers the logline's question, even ambiguously, and build your last two shots backward from it.
FAQ: AI Short Film Prompting
How long should a prompt be? Between 40 and 90 words for most models. Long enough to cover all six layers, short enough that the important details stay prominent.
Can I make a coherent film with text-to-video alone? Yes, for mood-driven pieces and landscapes. For anything with a recurring character, combine still generation with image-to-video to keep identity stable.
How do I keep a character consistent across shots? Write a character block and reuse it verbatim, generate a strong reference still, and drive later shots from that still. Wardrobe and prop anchors help the viewer track continuity even when the face shifts slightly.
How many takes should I generate per shot? Three to five variations, then stop. If none work, rewrite the prompt rather than generating more.
What aspect ratio should I use? Match your destination. Vertical for social, 16:9 for short film festivals and YouTube, and choose early because reframing later costs quality.
Do I need editing software? Yes. Any editor that handles multiple tracks and basic color correction is enough. The edit is where separate clips become a film.
Should I write dialogue in prompts? Only with a model that supports speech. Otherwise, keep characters silent, use voiceover recorded separately, and avoid mouth-visible shots during speech.
What is the single biggest beginner mistake? Generating before planning. Ten minutes with a shot list saves hours of scrolling through clips that do not cut together.
How do I handle complex action scenes? Break them into more, shorter shots. Fast cuts hide the limitations of individual generations and read as intentional energy rather than approximation.
When should I use live action instead? When a performance, a specific face, or physical interaction is central to the story. Hybrid approaches work well: AI for establishing shots and transitions, live footage for the emotional core.


