Cinematography used to be the most intimidating part of filmmaking. It required knowing which lens to use, where the light should fall, how to move the camera, and why any of it matters. AI video tools have not replaced that knowledge; they have made it available to anyone who can describe it. The same rules that make a movie scene look expensive apply when you are writing prompts.
The creators producing the best AI footage are not the ones with secret prompting tricks. They are the ones who understand a handful of cinematography principles and express them clearly. This guide covers those principles, from composition to camera movement to consistency, and shows how to apply them to AI generation.
Shot Composition: What to Describe First
Composition is where you decide what the audience looks at and what they feel about it. The three tools that matter most are framing, rule of thirds, and depth.
Framing decides how close the audience is to the subject. A wide shot establishes the world, a medium shot carries dialogue and action, and a close-up isolates emotion. Describe the framing explicitly: "wide shot of a city street at dawn" and "close-up of a character's eyes" are different movies.
The rule of thirds is a simple mental grid. Place the subject off-center, at one of the intersections of the grid, and the frame feels dynamic. Place it dead center and the frame feels confrontational or monumental. Both are useful; the mistake is choosing the center by default.
Depth is the third tool. Foreground, subject, and background layers make a frame feel real. Describe layers in your prompt: "a character in the foreground, a marketplace behind, mountains in the distance." Flat prompts produce flat images.
Camera Movement: Directing Motion in Text
Camera movement is the quickest way to make AI footage feel cinematic, because movement is what video does that photos cannot. The vocabulary is small and precise:
- Push-in: the camera moves toward the subject, increasing tension or intimacy;
- Pull-back: the camera moves away, revealing context and reducing pressure;
- Pan: the camera turns horizontally, revealing space;
- Tilt: the camera turns vertically, revealing height or scale;
- Tracking: the camera moves alongside the subject, creating energy;
- Crane or aerial: the camera rises or flies, establishing geography.
Write the movement into the prompt as an instruction: "slow push-in on a character reading a letter" or "aerial shot revealing the full valley." If the prompt says nothing about movement, the model invents it, and invented movement is usually the safest, most boring option. You are the director; specify the shot.
Depth of Field and Focus as Storytelling Tools
Focus tells the audience what matters. A shallow depth of field blurs the background and isolates the subject, which is the workhorse of emotional close-ups. A deep focus keeps everything sharp, which suits landscapes and group scenes. And focus can move: rack focus, shifting from one subject to another in the same frame, directs the eye mid-shot.
In a prompt, describe both the depth and the focus behavior:
- "shallow depth of field, background bokeh" for isolation;
- "deep focus, everything sharp from foreground to horizon" for scale;
- "focus shifts from the bottle in the foreground to the character behind the bar" for storytelling.
Many models handle rack focus surprisingly well when asked. The request just needs to be explicit, because an implicit request is a request the model will guess wrong.
Lighting and Mood Control
Light is the cheapest special effect in cinema, and it is fully available in AI prompts. The language of lighting is simple: source, direction, quality, and color.
Source is where the light comes from: window light, neon signs, candlelight, sunlight, or practicals in the frame. Direction tells the model where shadows fall: front light flattens, side light sculpts, backlight separates the subject from the background. Quality describes hardness: direct sun creates hard shadows; overcast creates soft, flattering light. Color sets emotion: warm for comfort, cold for tension, green for unease.
A complete lighting sentence in a prompt: "moody side lighting from a neon sign, hard shadows, cold blue tones." The model will build the whole scene around that sentence, and the mood will land.
Consistency Across a Sequence
A beautiful single shot is easy. A beautiful sequence is a discipline problem. The cinematography choices you make for one shot have to survive the transition to the next, or the film falls apart.
The practical rules are the same ones used on real sets:
- Lock the look: decide the color palette and lighting direction once, and repeat them in every prompt;
- Anchor the subjects: reference images and fixed descriptions for every recurring character and location;
- Match the grade: do a color pass at the end so every shot lives in the same world;
- Keep the grammar consistent: if the story is told with push-ins and close-ups, do not randomly cut to a crane shot without a reason.
Consistency is not monotony. It is a shared visual language that lets you break the pattern deliberately when the story needs a shock.
Choosing a Model for Photorealism vs. Style
The same cinematography principles produce very different results depending on the model. Match the model to the look you are chasing:
- Photorealistic footage: choose a model with strong physics and skin rendering; realism is the hardest task and needs the best tool;
- Stylized or animated looks: pick a model that understands art direction; style models punish hyper-realistic prompting;
- Fast drafts: use a quick model to block out shots, then move the keepers to the final model.
When in doubt, generate the same shot on two models and compare. The comparison teaches you more than any review article, and it takes minutes.
A Practical Capture Checklist
Before you generate, run this checklist so the shot has a fighting chance:
- Framing: wide, medium, close-up, or extreme close-up, and why;
- Subject position: center or off-center, and what that communicates;
- Layers: foreground, subject, background;
- Camera movement: static, push-in, pull-back, pan, tilt, tracking, or aerial;
- Focus: shallow, deep, or a rack focus;
- Lighting: source, direction, quality, and color;
- Mood: the single emotion the frame should produce;
- Consistency: same references and look as the rest of the sequence.
The checklist takes two minutes per shot and saves hours of regeneration. The shots that fail are almost always the ones where the checklist was skipped.
Three Cinematic Templates You Can Steal
Templates are a fast way to learn by example. Three reliable looks:
- The moody interior: "medium close-up, shallow depth of field, side window light, hard shadows, muted green and brown tones, slow push-in." Works for drama, interviews, and product storytelling;
- The golden hour exterior: "wide shot, warm backlight, soft lens flare, dust in the air, gentle tracking along the subject, amber and peach tones." Works for romance, adventure, and aspirational brand content;
- The neon night: "medium shot, cyan and magenta neon, wet street reflections, rain, rack focus from the sign to the character, high contrast." Works for thrillers, music, and urban style content.
Keep these templates in your notes, then modify the subject, the action, and the mood to make them yours. Templates give you a reliable starting point without making every video look the same.
Prompt Patterns for Common Shot Types
Different shots need different prompt shapes. The establishing shot asks for geography: "wide aerial shot, city skyline at dawn, mist, low sun, slow lateral movement." The two-shot asks for relationship: "medium two-shot, two characters facing each other across a table, shallow depth of field, warm practical light, camera drifting slowly." The insert shot asks for detail: "extreme close-up, a hand pouring coffee, steam rising, soft light, shallow focus, subtle handheld motion." The point-of-view shot asks for immersion: "POV shot, walking through a doorway into a bright room, lens flare, slight camera sway." One sentence of purpose plus one sentence of camera and light covers most shots.
Aspect Ratios and Formats
Format is a storytelling choice. Sixteen-by-nine is the default for web and long-form. Nine-by-sixteen suits vertical platforms and isolates subjects in tall frames. A 2.39:1 cinematic bar signals "film" instantly and flatters wide compositions, though you may need to generate wider than you crop. Changing format after generation wastes shots, so decide the ratio when you write the shot list. If a crop is required, regenerate the affected shots in the target ratio instead of stretching, which destroys the look you worked to build.
A Starter Shot Library
A small library of reusable shot recipes speeds up every project. Keep these patterns in your notes and adapt them:
- The arrival: "wide shot, a figure enters from the left of frame, camera pans to follow, golden hour light" for entrances and first appearances;
- The reveal: "medium shot, subject turns toward camera, slow pull-back, soft focus to sharp" for introductions and plot turns;
- The chase: "tracking shot, subject runs toward camera, handheld shake, fast pace, harsh daylight" for energy and urgency;
- The quiet moment: "static close-up, subject breathing slowly, window light, shallow focus, no movement" for emotion and reflection;
- The transition: "extreme close-up on an object, then the camera tilts up to reveal the new scene, matching color tones" for time or location changes;
- The ending: "wide establishing shot, subject small in frame, slow push-out, fading light" for closure and scale.
Recipes are starting points, not formulas. Change the subject, the light, and the mood, and the same skeleton serves a completely different story.
FAQ
Q: Do I need to study film to write good cinematography prompts?
A: A little vocabulary goes a long way. The terms in this article cover most of what you will use.
Q: Why do my generated shots look flat?
A: Usually missing layers or lighting. Add a foreground and background, and give the light a direction and color.
Q: Can AI do complex camera moves like a crane shot?
A: Some models handle them, but complex moves are the least reliable. When a move fails, simplify: a push-in or pan is more reliable and often reads just as well.
Q: How important is the model for cinematic quality?
A: Very. But a good prompt on a good model beats a great prompt on a weak model, and a great prompt on a good model beats both.
Q: Should every shot have camera movement?
A: No. Static shots have their own power, especially for tension. Movement should be a choice, not a default.
Q: How do I know if my prompt is too long?
A: If the model ignores half of it, it is too long. Cut the adjectives, keep the structure: framing, subject, action, camera, light, mood.
Q: Can AI handle slow motion?
A: Sometimes, but it is unreliable. If slow motion is essential, generate at normal speed with clear action and slow the clip in post; the result is usually smoother than asking the model.
Q: Should I always describe camera movement in the prompt?
A: Yes, when movement matters. A static prompt leaves the camera to chance; describing the move turns the shot into a directorial choice.
Q: How do I match generated shots with real footage?
A: Match the light direction, color grade, and lens feel in the prompt, then do a final grade across the whole sequence so the seam disappears.
Q: What is the single highest-impact skill for AI cinematography?
A: Writing the camera and the light before the subject. Most prompts describe what is in the frame; the shots that look directed describe how the frame is seen.
Final Thoughts
AI cinematography is not a new art form so much as the old art form with a new instrument. The principles, composition, movement, focus, light, and consistency, all transfer directly. What changed is that you no longer need a crew to execute the shot; you need to know what to ask for. Build the vocabulary, apply the checklist, and your footage will look less like random generation and more like a film someone actually directed.



