Vente à Durée Limitée : Profitez de 30% DE RÉDUCTION sur la Création Vidéo IA de Nouvelle Génération 🎉

AI-Assisted Cinematography: Lighting and Composition Secrets

Sep 15, 2026

Why Light and Framing Still Decide Whether AI Video Feels Cinematic

Most viewers cannot name a camera model or a codec, but they can feel the difference between a clip that looks like a home video and one that looks like a scene from a film. That feeling comes from two things: the arrangement of elements inside the frame, and the way light sculpts those elements. Everything else — resolution, motion smoothness, grading polish — amplifies those two foundations, but never replaces them.

This matters more than ever now that generative video is genuinely accessible. A single prompt can produce stunning textures in seconds, yet that same prompt will produce shots that feel flat and stagey if it says nothing about where the camera stands or how the subject is lit. The craft did not disappear. It moved. Instead of positioning lights on a physical set, you now describe them. Instead of dragging a tripod across the floor, you state a lens, a height, and a distance.

This guide is written for people who want control rather than luck: solo creators, small teams, marketers producing narrative spots, and filmmakers prototyping sequences before a shoot. It covers how to think about lighting and composition, how to write prompts and shot plans that actually constrain them, how to use references without destroying character consistency, and how to build a repeatable workflow that keeps a project visually coherent from the first shot to the last.

The Two Languages of the Frame

Composition decides where attention goes

Composition is attention management. Every choice — subject placement, negative space, foreground obstruction, the direction of a gaze — either narrows the viewer's focus or scatters it. Classic tools like the rule of thirds, leading lines, and frames-within-frames are not laws handed down from cinema history; they are shortcuts for answering one question: what should the eye do next?

In AI video work, composition doubles as a consistency device. If your protagonist is always framed slightly left of center with their gaze leading into the empty right side of the frame, cuts between shots feel coherent even when the backgrounds are wildly different. If every shot invents its own framing logic, the edit will feel random no matter how beautiful the individual clips are.

Lighting decides what the frame means

Lighting carries the emotional load. Hard, directional light from a single source creates tension, secrecy, or menace. Soft, broad light creates safety, intimacy, or nostalgia. Backlight separates a subject from the background and manufactures depth. Practical sources — lamps, screens, neon signs — anchor a scene in a believable world and give you motivated color for free.

The practical takeaway: decide the emotion first, then choose light quality, direction, color temperature, and contrast ratio to serve it. Beginners usually do the opposite. They write "cinematic lighting" into a prompt and hope an emotion appears on its own. It rarely does.

Where the two systems meet

Light and framing are one system, not two. A hard key from camera left drops a shadow across the face, which changes where the viewer's eye lands inside the composition. A wide frame with a small subject needs either a bright pool of light to localize attention or a strong geometric shape to hold the eye. When you plan shots, describe both in the same sentence, because that is how they function on screen.

Reading a Scene Before You Generate It

Before touching any tool, write a shot card for every shot you need. A shot card is deliberately small:

  • Subject and action
  • Emotional beat
  • Shot size (wide, medium, close)
  • Camera height and angle
  • Lens feel (wide, normal, telephoto)
  • Light quality, direction, and color
  • One dominant background element
  • Movement (static, push in, orbit, handheld)

Eight lines per shot is enough, and the discipline of filling them in forces clarity. The reason this works is that generative models respond to concrete constraints far better than to adjectives. "Cinematic" tells a model nothing useful. "Medium shot, eye level, 50mm feel, single soft key from the left, warm tungsten practical in the background" tells it almost everything.

Keep the shot cards in a plain text file or spreadsheet, one row per shot. When you need to regenerate, extend, or replace a shot weeks later, the card is your memory. It also makes collaboration possible: a writer, an editor, and a model operator can all work from the same grid without spending an hour re-deciding what the scene looked like.

Lighting Setups You Can Describe in Plain Language

You do not need a lighting diagram. You need a vocabulary that maps reliably to generated output. These five setups cover the vast majority of narrative needs.

Three-point lighting

Key, fill, and backlight. The key defines the shape of the face, the fill controls how dark the shadows go, and the backlight separates the subject from the background.

Prompt language: "soft key light from camera left at 45 degrees, gentle fill from camera right, rim light behind the subject."

Best for: dialogue, interviews, character introductions. It is the safest, most legible look and the easiest to keep consistent across a series of shots.

Single-source and Rembrandt-style light

One hard source, deep shadows, a small triangle of light on the shadow-side cheek. High contrast, high drama, minimal information.

Prompt language: "single hard light from upper left, deep unlit shadows, high contrast, chiaroscuro."

Best for: tension, confrontation, mystery, characters who are hiding something.

Silhouette and backlight

The subject is dark against a bright background. The audience reads shape and posture instead of expression.

Prompt language: "strong backlight through a doorway, subject in silhouette, lens flare, dust visible in the air."

Best for: entrances, reveals, endings, and any moment when a character should feel inaccessible.

Practical-motivated light

The source is visible in the frame: a lamp, a fire, a screen, a car headlight. Color and direction come from that object rather than from an invisible rig.

Prompt language: "lit only by a desk lamp in the lower frame, warm pool of light, the rest of the room falling into darkness."

Best for: night interiors, intimacy, and scenes where realism matters more than glamour.

Ambient and available light

Overcast sky, window light, shaded outdoor spaces. Soft, low contrast, natural color, gentle shadow transitions.

Prompt language: "soft overcast daylight through a large window, low contrast, cool neutral tones."

Best for: quiet scenes, documentary texture, and projects that need to feel unforced.

Two dials worth memorizing

Beyond the setup names, two controls do most of the emotional work. The first is color temperature: warm light (amber, orange, candle) reads as safe, nostalgic, or intimate, while cool light (blue, cyan, moonlight) reads as clinical, lonely, or threatening. Mixing warm and cool in one frame — a warm face against a cool room — instantly creates separation and depth.

The second is contrast ratio, the relationship between the brightest and darkest parts of the frame. A low ratio feels gentle and documentary. A high ratio feels theatrical and tense. When a shot feels flat, the ratio is usually the culprit, not the color palette.

Composition Rules Worth Keeping and the Ones Worth Breaking

Keep: the rule of thirds as a default

Placing a subject on a third line gives the frame tension and leaves room for movement in the direction of the gaze. It rarely looks wrong, and it gives you predictable framing you can reuse across an entire scene.

Keep: leading lines and depth layers

Backgrounds with converging lines — roads, corridors, shelves, fences — pull the eye toward the subject. Layering foreground, midground, and background creates depth, and depth is what separates cinematic imagery from flat illustration. Even a single out-of-focus foreground element can transform a mediocre frame.

Keep: headroom and eyeline discipline

Consistent headroom is one of the strongest continuity signals in an edit. If shot one has generous headroom and shot two crops the forehead, the cut feels amateurish even when both frames are gorgeous on their own. Decide the headroom convention once, write it down, and apply it to every shot.

Keep: negative space with intent

Empty space is not wasted space. A figure placed in the corner of a wide frame with nothing but sky or wall beside them communicates isolation before a single line of dialogue lands. Just make sure the emptiness is doing something specific, because lazy empty space just looks like bad framing.

Break: symmetry when you want formality or unease

Centered framing is powerful for confrontation, ritual, and control. It can also signal that something is slightly off. That makes it a good sparing choice in a scene that otherwise uses off-center framing — but a poor default, because constant symmetry flattens the rhythm.

Break: rules when the story demands discomfort

Dutch angles, extreme wide shots with tiny subjects, and cramped framing communicate disorientation, isolation, and pressure better than any lighting change. Use them deliberately and rarely so they keep their force. A tilted frame that appears in every scene stops meaning anything.

Aspect ratio is a composition decision

Widescreen emphasizes environment and horizontal movement; a squarer frame pushes faces and bodies forward and makes enclosed spaces feel tighter. Pick one ratio for the project and stay there, except when a deliberate format change marks a shift in time, memory, or perspective.

Writing Prompts That Actually Control Light and Framing

A reliable prompt has five parts, in this order:

  1. Shot size and angle — "low-angle medium close-up"
  2. Subject and action — "a woman in a rain-soaked coat looks up"
  3. Light: quality, direction, color — "hard cool light from behind, warm practical spill on her face"
  4. Composition and lens — "subject on the right third, shallow depth of field, 85mm feel"
  5. Mood and texture — "tense, wet pavement reflections, fine grain"

Two rules make this work. First, avoid contradictory light directions. Asking for a soft key from the left and hard light from the right produces mush, because the model averages the conflict. Second, place composition constraints near the end of the prompt, where they act as a final instruction rather than competing with the subject description at the front.

Negative instructions are much weaker than positive ones. Instead of "no harsh shadows," write "soft shadows with visible detail in the dark side of the face." Models handle describable states better than absences, so every "don't" is a missed opportunity to specify a "do."

Finally, vary one variable at a time when you are testing. If you change the lens, the light direction, and the framing in a single iteration, you learn nothing about which change produced the improvement.

Keeping a Project Visually Consistent

Consistency in AI video is a documentation problem more than a model problem. Four habits carry most of the weight.

Lock the palette. Choose three dominant colors and one accent, then state them in every prompt. A project shot in teal, amber, and warm gray will feel unified even if individual shots vary widely.

Lock the lens language. Pick one primary focal-length feel and one secondary. Mixing extreme wide and extreme telephoto inside one scene is a stylistic statement, not a neutral default.

Lock the light logic. Decide whether your world is motivated by practicals, daylight, or stylized keys — and stay there. Audiences forgive repetition far more readily than inconsistency.

Lock the aspect ratio and grain. Changing ratio mid-scene is jarring, and changing grain structure between shots reads as a quality drop even when nothing else changed.

Write these four locks at the top of your shot file. Before you generate, check the card against the locks. This single habit eliminates most "why does this shot feel wrong" problems.

Reference images and style transfer without losing your subject

Reference-driven workflows are the fastest way to raise visual quality, and also the fastest way to lose character consistency. Treat references as two separate inputs. Look references carry lighting, palette, grain, and lens character. Identity references carry face, wardrobe, and distinctive props.

When you blend them, describe the split explicitly: "match the lighting and color palette of the reference; keep the subject's face and clothing from the character sheet." If a tool offers separate slots for style and subject, use them — that separation is the entire point of the feature.

Keep a small, curated look library of three to five images per project. Too many references produce averaged, generic results. Record which reference influenced which shot, because you will need to reapply the same look when you add or replace shots later.

A Practical End-to-End Workflow

Step 1: Script and break down. Convert the script into a shot list of ten to thirty shots, each with a completed shot card. Do this before generating anything, and resist the urge to skip it because the tool is fast.

Step 2: Build a look book. Gather three to five lighting references and one character sheet, then write the four locks.

Step 3: Generate still frames first. Iterate on composition and light at the image level, where each attempt is cheap and quick. Approve a still for every shot before animating anything. This one reordering saves more time than any other change you can make.

Step 4: Animate the approved stills. Convert to video, extend duration, and keep camera movement minimal unless the shot genuinely needs it. Movement is where most consistency failures originate, because every frame is a new opportunity for the subject to drift.

Step 5: Assemble a rough cut early. Edit with placeholder timing rather than waiting for perfect clips. You will discover missing coverage — inserts, reaction shots, establishing frames — before you have spent effort generating material you do not need.

Step 6: Fill gaps with inserts. Close-ups of hands, objects, and environment shots are inexpensive to produce and solve most continuity problems in the edit.

Step 7: Unify in post. Apply one color pass and one grain pass across the whole timeline. A single look applied to every clip masks small inconsistencies between generations.

Step 8: Review with sound on. Music and sound design change how lighting and framing read. A slow push-in that felt sluggish in silence may feel exactly right under a rising score.

On time budgeting: stills typically consume a third of your production time, video generation another third, and assembly plus grading the rest. If animation is eating half your schedule, the problem is usually unresolved stills, not slow rendering.

Common Mistakes and How to Fix Them

Everything is "cinematic" and nothing is specified. Fix: replace adjectives with physical descriptions of light and camera.

Every shot is a medium shot. Fix: alternate shot sizes deliberately — wide, medium, close — so the edit has rhythm.

Lighting direction changes between shots in the same scene. Fix: state light direction on the shot card and copy it forward verbatim.

Subjects are always centered with identical headroom. Fix: apply the rule of thirds by default and reserve centered framing for emphasis.

Too many references. Fix: cut the look library to three to five images, each with a written reason for inclusion.

Movement everywhere. Fix: default to static or slow movement and save fast moves for moments that need energy.

No coverage for the edit. Fix: generate inserts and reaction shots before assembly, not after a failed first cut.

Grading every clip individually. Fix: grade once, across the timeline, in a single pass.

FAQ

How much of cinematography can AI actually handle? Generation tools handle rendering and will follow detailed instructions about light and framing, but the decisions — what a shot means, where the camera goes, how shots cut together — remain yours. AI removes execution cost, not authorship.

Do I need real film knowledge to get good results? You need vocabulary, not experience. Learning five lighting setups and three composition patterns is enough to dramatically improve output quality.

Which matters more, lighting or composition? Composition, if you can only fix one. A well-composed shot with mediocre light still reads clearly; a badly composed shot with beautiful light usually just looks confusing.

How do I keep a character consistent across many shots? Use a character sheet, repeat a fixed identity description in every prompt, and change only the shot-specific variables around it.

Why do my shots look flat? Usually because light is coming from the camera position and contrast is low. Move the key off-axis, add a rim light, and increase the ratio between lit and unlit areas.

Should I generate stills first or go straight to video? Stills first. Iterating on composition and lighting is faster and cheaper at the image stage, and approved stills make animated results far more predictable.

How many shots should a short scene have? For a one-minute scene, eight to fifteen shots is a workable range, including at least two inserts and one wide establishing frame.

What is the fastest way to improve? Copy one reference frame you love, write down its light direction, contrast, and subject placement, then rebuild that shot with your own subject. Repeat with five different references. That exercise teaches more than any preset collection, because it builds the habit of translating images into decisions you can control.

Alexander

Alexander