Anyone who writes prompts for AI video has felt the frustration of the near-miss. The scene is fine, the motion is plausible, but the footage still looks like generic AI slop rather than a deliberate piece of filmmaking. The missing ingredient is usually not better technology but better visual literacy: a real vocabulary for what a shot is, why certain shots feel good, and how camera choice communicates story. That vocabulary was written down decades ago by two film scholars, David Bordwell and Kristin Thompson, and it turns out to be exactly what modern prompt writers need. This article decodes their concept of the shot and shows how to turn it into practical, frame-by-frame control over AI-generated video.
Why a 1980s film textbook is the best prompt-writing manual
Bordwell and Thompson wrote extensively about how films are put together, and their framework has aged better than almost anything in media studies. Their core idea is that film style is the purposeful selection and arrangement of techniques, not random accumulation. Every choice a filmmaker makes about framing, camera movement, lighting, and editing is a decision about what an audience should feel and understand at that moment.
When you generate a video from a text prompt, you are effectively writing a one-sentence directing note for an automatic cinematographer. If you cannot think at the level of the shot, your note will be vague, and a vague note produces vague footage. The fix is to internalize the shot as the basic unit of storytelling and to let that unit structure every prompt you write.
What a shot actually is, and what it is not
It is easy to confuse a shot with a scene. Bordwell and Thompson are careful to separate them. A scene is made up of a series of shots, and a shot is a single continuous run of the camera, an uninterrupted sequence of frames with no edit in the middle.
Seen that way, the shot is the smallest unit of visible story material. Its duration, framing, angle, and internal movement all contribute meaning. When you prompt for AI video, you are asking for exactly one shot at a time, and naming it as a shot changes how you write. Instead of describing a whole scene and hoping the generator resolves it, you describe a single continuous take, one that respects the logic of a real camera.
This distinction matters in practice. A prompt like a knight walking through a forest reads as a scene request and tends to produce a montage of disconnected imagery compressed into a few seconds. A prompt that says a low-angle tracking shot follows a knight as he walks forward through misty forest, one continuous take keeps the camera glued to a single event and gives the generator a clear unit to deliver.
The basic building blocks: framing, angle, and distance
Bordwell and Thompson organize shot analysis around a handful of formal elements, and the first cluster is framing. Framing is how the camera frames the subject, and it has three useful components.
Distance is the most obvious: extreme long shots establish place, long shots show the body and context, medium shots favor the waist up and dialogue, and close-ups prioritize faces and small details. Each distance is a storytelling choice. Close-ups manufacture intimacy and significance; long shots manufacture isolation or scale. When you prompt, pick a distance on purpose rather than letting the generator guess.
Angle is the vertical relationship between camera and subject. A high angle looks down on a subject and can make it feel small, vulnerable, or watched. A low angle looks up and can make a subject feel powerful or imposing. A neutral eye-level angle is the honest, default documentary register. Angle is a cheap way to inject a lot of meaning into a prompt.
Height and level follow. Height is where the camera sits relative to the ground, and level is whether the frame horizon is tilted. A tilted or canted frame signals unease, instability, or a disrupted worldview. Only include such choices when the story wants them. Describing these elements explicitly is the difference between a generic image and a directed one.
Camera movement as narrative instrument
Bordwell and Thompson give camera movement its own analytic weight because movement changes what we see and how we interpret it. A pan rotates the camera horizontally from a fixed position and reveals space. A tilt does the same vertically. A tracking shot physically moves the camera alongside the action, and a crane shot changes height mid-take.
The crucial insight is that movement does not just show motion; it guides meaning. A tracking shot that follows a character implies we should stay with them, that they matter. A handheld shake implies immediacy, documentary truth, nervousness. A slow push-in toward a face increases tension or emphasis. A pull-back can reveal isolation or scale.
In prompt terms, describing camera movement precisely gives the generator a reliable instruction. Writing a slow dolly-in toward the subject's eyes is far more controllable than saying focus on the emotion. The former names a specific motion; the latter leaves cinematography to chance. You can also chain movement deliberately: an establishing crane shot that descends into a tracking follow signals a change from context to character.
Mise-en-scène: everything the camera frames
Bordwell and Thompson use the French term mise-en-scène, literally putting into the scene, to name everything placed in front of the camera: setting, costume, lighting, props, and blocking. A director controls meaning not only through the camera but through what the frame contains and how those elements are arranged.
For a prompt writer, mise-en-scène is the richness your description adds beyond the camera. Lighting direction and quality are huge. A hard key light with deep shadows reads as film noir; soft, warm, diffused light reads as gentle and nostalgic. The layout of characters in the frame, your blocking, tells the audience about relationships: who is centered, who stands apart, who dominates the space.
Treating mise-en-scène as a deliberate toolkit transforms prompts from thin sentences into dense director's notes. Instead of saying a rainy city street at night, you say a rain-slicked street lit by a single amber streetlight, a lone figure standing in shadow on the right, steam rising from a grate, and the generator responds with layered, composed footage rather than a generic puddle.
Depth of field as a focus for the eye
Bordwell and Thompson discuss depth of field as a way of organizing attention within a frame. A shallow depth of field keeps only a narrow slice sharp and throws the background into soft blur, isolating the subject and simplifying the composition. A deep depth of field keeps foreground and background sharp, inviting the eye to explore the whole scene.
Prompt writers can exploit this directly. Shallow depth of field is a strong tool for portraits and dialogue, where focus on the face matters more than environment. Deep focus suits scenes where the environment itself is the story, like a wide establishing shot of a landscape or a crowded space. Naming the depth of field in the prompt is an instant clarity and emotion lever.
The shot on its own is never enough: editing and montage
A single beautiful shot is not yet a film. Bordwell and Thompson insist that meaning forms across the cut: a montage or continuity sequence. Two shots placed side by side create a meaning neither holds alone, the so-called Kuleshov effect. The gap between shots is where narrative rhythm lives.
For AI video, this is the reminder that a strong single take is still only one frame of a larger argument. You should plan a sequence of shots the way a director plans a scene's coverage. Write each shot as a deliberate step, then assemble them in editing so the juxtaposition creates meaning. The best AI projects are shot lists, not isolated generations.
Continuity: keeping the world consistent across shots
Continuity is the invisible glue of a sequence. Bordwell and Thompson stress matching action, screen direction, and spatial relationships so the audience never gets lost. If a character walks screen left in one shot, they should generally continue screen left in the next, keeping the space coherent.
In AI workflow, continuity is the hard problem of keeping a character, location, and look identical across multiple takes. The shot-level mindset helps by making your references explicit before each generation: describe the same costume, the same lighting register, the same lens feel. Combine that with the reusable reference imagery and prompt templates discussed in the workflow section so each shot in the sequence feels like it comes from the same camera roll.
Sound as the neglected half of the shot
Bordwell and Thompson treat sound as a full partner to the image. The pairing of what we see and what we hear shapes meaning every bit as much as the visuals. Sound can ground a scene chronologically, create tension, or tell us what to feel before the image changes.
If your AI video pipeline allows audio, design it as carefully as the image. A shot of an empty corridor means very different things with soft wind versus a ticking clock versus distant laughter. Even the decision about silence is a decision. Matching the audio register to the shot, and letting the sound change with the cut, finishes what the image starts.
From theory to practical prompting: a worked example
To see the payoff, compare two prompts for the same idea on paper.
Vague draft: A detective enters a dark office and finds a clue.
That is a scene request, and it produces exactly the mush you would expect: a rapid montage of doors, rooms, papers, and faces mashed into a few seconds with no single clear take.
Director's version, thinking in shots: Shot one, a low steady dolly-in at waist height follows the detective as he pushes open a heavy oak door; a single cold overhead light and deep shadows suggest danger; shallow depth of field keeps him sharp against the blurred office. Result, one continuous, controlled take with clear tone.
The second version works because it names the shot, the framing, the angle, the movement, the lighting, and the depth of field. Each term is a controllable instruction the generator can execute, and together they add up to deliberate filmmaking.
Frequently asked questions
Do I need to study film theory to write good prompts? You need the vocabulary, and the core concepts here, shot, framing, angle, movement, mise-en-scène, depth of field, continuity cover most of it. Bordwell and Thompson give you a rigorous frame, but you can apply it immediately without a course.
How do I keep AI footage from jittering between shots? Enforce continuity in every prompt: repeat the same character description, costume, lighting, and lens register, and use consistent reference imagery. A shared shot list that carries these details across all takes is the practical fix.
Is every shot a single generation? Generally yes, generate shot by shot and assemble in post. Trying to generate a whole scene in one prompt usually collapses into incoherent montage. Sequence discipline is what separates film from collage.
What is the biggest mistake prompt writers make? Describing scenes, not shots. The moment you start thinking one continuous take at a time and naming distance, angle, movement, and light, the footage transforms.
Can these principles work for fast social edits? Absolutely. Even a 15-second reel is a sequence of shots. Applying shot logic to each cut makes short-form content read as intentional, which is exactly what stands out in a feed.
Final thoughts
The shot is where a film is truly directed, and Bordwell and Thompson gave the concept formal clarity decades before anyone dreamed of text-to-video. The technology has changed, but the language has not; it is simply migrating from the camera assistant's notes to the prompt field. When you stop describing whole scenes and start planning individual shots, naming framing, angle, movement, lighting, depth of field, and continuity, you stop writing lottery tickets and start directing. That is the difference between generated footage and a film.


