Limited Time Offer: Get 50% OFF your first month of Pro & Ultra plans 🎉

Footcandles to F-Stops: Cinematography for AI Video

Sep 15, 2026

Why Cinematography Vocabulary Still Matters in AI Video

Generative video tools have made one thing clear: you no longer need a camera crew to produce a moving image. You still need to know what you want the image to look like. The gap between a flat, vaguely plastic-looking clip and a shot that feels like it came off a real set is rarely about resolution or frame rate. It is almost always about light and focus.

Traditional cinematography describes those two forces with a compact vocabulary. Light falling on a surface is measured in footcandles or lux. The size of the lens opening is described as an f-stop, which in turn governs depth of field. Exposure is the negotiation between them, plus sensitivity and shutter. These terms did not disappear when diffusion models arrived; they became instructions. A well-crafted prompt that says "warm practical lamp at f/2, shallow focus on the foreground cup, background falling into soft bokeh" is doing the same job a director of photography does when calling for a specific stop and a specific lighting setup.

This guide is a practical translation layer. It explains what footcandles and f-stops actually mean, how they interact, and how to convert that understanding into repeatable prompt patterns, shot lists, and review criteria for AI-generated footage. The goal is not to turn you into a camera technician. The goal is to give you a shared language so your intent survives the trip from your head to the model.

Lighting Units, Decoded: Footcandles, Lux, and Real-World Reference Points

A footcandle is the amount of illumination produced by one candle at a distance of one foot. That definition sounds quaint, but the unit remains useful because it is intuitive and because decades of production knowledge are recorded in it. The metric equivalent, lux, is roughly 10.76 times larger: one footcandle equals about 10.76 lux.

Reference values worth memorizing

The numbers below are approximations, but they are close enough to guide decisions:

  • Moonlit exterior: about 0.01 to 0.1 footcandles
  • Candlelit interior: roughly 1 to 2 footcandles
  • Living room with practical lamps at night: 5 to 20 footcandles
  • Bright office interior: 30 to 50 footcandles
  • Overcast daylight exterior: 100 to 500 footcandles
  • Direct sun on a clear day: 1,000 to 10,000 footcandles

What matters for AI video is not the exact figure. What matters is the ratio. Cinematographers think in contrast ratios far more than in absolute levels. A face lit with a key at 100 footcandles and a fill at 25 footcandles gives a 4:1 ratio, which reads as moody but readable. A 1:1 ratio reads as flat and clinical. A 16:1 ratio reads as noir.

Turning footcandle thinking into prompt language

Models do not meter light. They pattern-match against descriptions and against the visual statistics of their training data. So instead of writing "light the scene at 40 footcandles," write what that level looks like:

  • "Soft window light filling the room, gentle shadows, mid-morning overcast quality"
  • "Single practical lamp, warm pool of light, deep falloff into the corners"
  • "Harsh midday sun through blinds, bright contrast stripes across the wall"

If you want to be more technical, you can include both: "interior lit to roughly 15 footcandles, warm 3200K practical, shallow falloff." Some models respond well to color temperature in Kelvin and to the words "practical" and "falloff." Even when a model ignores the precise number, the surrounding words steer the render toward the intended contrast pattern.

Why ratio beats level

The single most common lighting mistake in AI video is asking for a bright, evenly lit scene and then wondering why it looks like a stock photo. Even illumination removes the cues that make a frame feel dimensional. Push the ratio, keep a shadow side, and let the background sit two or three stops darker than the subject. That alone will make generated footage feel more photographed.

Aperture and F-Stops: Depth of Field as a Directable Variable

An f-stop is the ratio of a lens's focal length to the diameter of its effective aperture. Small numbers mean a large opening: f/1.4 lets in a lot of light and produces a very shallow plane of focus. Large numbers mean a small opening: f/16 lets in far less light but keeps most of the frame sharp.

What each stop does to your image

Each full stop changes exposure by a factor of two. The standard sequence is f/1.4, f/2, f/2.8, f/4, f/5.6, f/8, f/11, f/16, f/22. Going from f/2.8 to f/4 halves the light reaching the sensor. Going the other direction doubles it.

For AI video, the practical meaning is depth of field:

  • f/1.2 to f/2: very shallow. A face is sharp from nose to ear, everything else dissolves. Great for intimate portraits, dreamy sequences, and isolating a subject in a busy environment.
  • f/2.8 to f/4: the classic narrative range. Subject is crisp, background is soft but still legible.
  • f/5.6 to f/8: balanced. Both subject and immediate surroundings read clearly.
  • f/11 to f/16: deep focus. Landscapes, wide establishing shots, architecture, group scenes where everyone must be sharp.

Why f-numbers behave differently across formats

The same f-stop produces different depth of field on different sensor sizes. A full-frame sensor at f/2.8 has noticeably shallower depth of field than a Micro Four Thirds sensor at the same stop. Because generators do not have a sensor, you have to describe the format if the look depends on it. Phrases like "large-format look, f/2, creamy separation" or "16mm documentary texture, deeper focus" nudge the model toward the right optical character.

Bokeh is not just blur

Bokeh describes the character of out-of-focus highlights. Wide apertures with many-bladed apertures produce round, soft circles; older or simpler optics produce busier, more angular shapes. You can request this directly: "round, soft bokeh highlights on distant string lights" or "busy, swirly out-of-focus background." These small optical details are often what separates a convincing shot from an obvious render.

The Exposure Triangle in a Generator: Sensitivity, Shutter, and Gain

Real cameras balance aperture, shutter speed, and ISO. Generators do not have these controls, but they do have visual consequences that you can request.

  • High sensitivity (ISO 1600–6400): visible grain, slight color noise, reduced contrast. Request it for night exteriors, concert footage, and documentary realism.
  • Low sensitivity (ISO 100–200): clean, saturated, high dynamic range. Request it for commercial work and bright daylight.
  • Long shutter or motion blur: "180-degree shutter, natural motion blur on the hands" produces a filmic cadence. Too little blur makes motion look like a video game cutscene.
  • High frame rate look: "smooth 60fps slow-motion, water droplets suspended" for sports or action inserts.

A useful habit is to write a small "camera card" at the top of every prompt: format, lens, stop, sensitivity, and shutter feel. It takes ten seconds and it removes an enormous amount of randomness.

Building a Prompt Vocabulary That Maps to Camera Physics

Once you know the underlying concepts, the prompt becomes a compact technical brief. Think of it as four stacked layers.

Layer 1: Light source and quality

Name the source (window, practical lamp, neon sign, fire, bounce card), its direction (key from camera left, rim from behind), its quality (hard, soft, diffused), and its color (3200K tungsten, 5600K daylight, magenta neon). Quality and direction do more work than intensity.

Layer 2: Exposure and contrast

Describe the ratio, not the absolute level: "bright key, deep shadow side, background three stops under." Add a hint about highlight behavior: "retained highlights on the lamp shade, no clipping."

Layer 3: Optics

Focal length implies framing and compression. A 24mm lens exaggerates space; an 85mm lens compresses it and flatters faces. Combine focal length with an f-stop for a complete optical instruction: "85mm at f/1.8, tight portrait framing, background city lights as soft discs."

Layer 4: Motion and finish

Camera movement (slow dolly in, handheld drift, locked-off tripod), shutter feel, and grade (teal shadows, warm highlights, low contrast filmic curve). This layer is where most clips either feel cinematic or feel synthetic.

How Different Model Families Respond to Optical Instructions

Not every generator treats physics vocabulary the same way. Broadly, there are three behavioral groups, and knowing which one you are working with saves hours.

Literalist models tend to follow explicit camera language closely. They respond well to focal length, f-stop, and lighting-direction phrases, and they reward detailed briefs. If your prompt already reads like a shot card, these are your workhorses for controlled narrative work.

Aesthetic models prioritize overall look and mood over literal parameters. They may ignore "f/1.4" but react strongly to "dreamy, shallow focus, glowing highlights." With these, translate physics into mood adjectives and keep the technical layer short.

Motion-focused models excel at movement and physics but can drift on lighting continuity. Give them clear subject action and camera choreography, then fix lighting consistency in post or by reusing a consistent opening frame.

The practical takeaway: keep two versions of every prompt, one technical and one descriptive, and test both. Over a handful of generations you will learn which style your chosen model prefers.

A Practical Workflow: From Shot List to Finished Clip

Step 1: Write a shot list with exposure intent

Before opening any tool, write each shot in one line that includes subject, light, and optics. For example: "Close-up, hands pouring coffee, single warm practical from the left, 50mm at f/2, background kitchen falling dark." This forces you to decide the look before the model decides it for you.

Step 2: Collect references

Gather two or three still images that match your intended contrast and depth of field. You do not need to upload them, though many tools support image conditioning. Simply having them in front of you keeps your adjectives honest.

Step 3: Generate variants with one variable at a time

Change only the light ratio, then only the stop, then only the lens. If you change everything at once, you will never learn what the model actually responds to. Save the prompts that work into a personal library organized by look: "night interior practical," "overcast exterior deep focus," "neon rim portrait."

Step 4: Protect continuity across shots

Shots in the same scene must share a light direction, a color temperature, and a depth-of-field feel. Write these three values at the top of your project notes and paste them into every prompt for that scene. Inconsistency between cuts is the fastest way to make an AI sequence feel assembled rather than directed.

Step 5: Grade and finish

Generation is the middle of the process, not the end. Apply a light grade to unify color, add subtle grain to hide synthetic smoothness, and check that highlights do not clip and shadows do not crush. A tiny amount of vignetting can also help sell the lens.

Common Mistakes and How to Fix Them

Everything is sharp. You asked for a wide shot with no optical information, so the model defaulted to deep focus. Add an f-stop and a focal length.

Everything is flat. Your lighting description had no direction or ratio. Name a key direction and specify that the background falls off.

The subject looks pasted in. Usually a lighting mismatch: foreground and background share the same intensity and color. Push the background darker and cooler, or warmer if the scene is daylight-motivated.

Motion looks like a slideshow. Missing motion blur language. Request a 180-degree shutter feel and describe the action in terms of speed.

Faces look overly smooth. Ask for natural skin texture, visible pores, and a slightly imperfect key that leaves one side of the face in shadow.

Shot-to-shot drift. No scene-level camera card. Lock light direction, color temperature, and stop, then reuse them verbatim.

Decision Guide: When to Specify Physics and When to Describe Mood

Situation Best approach
Narrative dialogue scene Full technical brief: lens, stop, key direction, ratio
Product or food beauty shot Technical brief plus texture and highlight language
Abstract or dream sequence Mood-first language, minimal numbers
Documentary realism Available-light language, higher sensitivity, slight grain
Landscape establishing shot Deep focus, wide lens, time-of-day specificity
Action insert Shutter feel, motion blur, tight lens, short duration

A simple rule: the more the shot depends on dimension and separation, the more physics you should specify. The more it depends on atmosphere, the more you should lean on descriptive light language.

FAQ

Do I need to know math to use footcandles?
No. The unit matters less than the habit of thinking in levels and ratios. If you can say "bright key, dark background," you already have the useful part.

Should I put f-stop numbers directly in prompts?
Sometimes. Literalist models often respond well. Aesthetic models may ignore them, so pair the number with a visual consequence: "f/1.8, background lights as soft discs."

How do I get consistent depth of field across a series of shots?
Write one optical line per scene and reuse it word for word. Consistency comes from repetition, not from rephrasing.

What is the fastest way to make generated video look more filmed?
Add shadow. Uneven lighting with a clear direction and a darkened background instantly reads as photographed rather than rendered.

Is a shallow depth of field always better?
No. Shallow focus isolates but it also hides context. Use deep focus when the environment carries story information, and save wide apertures for emotional or product moments.

How much should I rely on post-production?
Treat generation as principal photography and post as finishing. A consistent grade, controlled grain, and unified contrast can rescue a sequence that was generated with slightly mismatched lighting.

The underlying principle is simple. Footcandles and f-stops are not trivia; they are the levers that control dimension, mood, and attention. Learn how they interact, translate them into concrete visual language, and apply them consistently across a project. The models will still surprise you, but the surprises will start landing inside the frame you intended.

Alexander

Alexander