Vente à Durée Limitée : Profitez de 30% DE RÉDUCTION sur la Création Vidéo IA de Nouvelle Génération 🎉

Prompt Engineering for AI Video Scripts: A Practical Guide

Sep 14, 2026

Why prompt engineering has become script engineering

For decades, the most expensive part of making a video was the edit. You shot too much footage, then spent weeks in a timeline trimming, colour grading, and fixing continuity problems that should never have reached the camera. Generative video changed the order of operations. When a model can produce a usable eight-second shot from a paragraph of text, the bottleneck moves upstream — into the written description that tells the model what to build.

That is why prompt writing now behaves like screenwriting. A prompt is no longer a search query. It is a shot description, a lighting plan, a wardrobe note, and a camera instruction compressed into a single block of language. The people producing the most reliable AI video are not the ones with the fanciest tools; they are the ones whose written plans are specific enough that a model has very little left to guess.

This guide is a practical workflow. It covers how modern text-to-video systems parse a prompt, how to convert a screenplay into shot-level instructions, how to keep a character recognisable across twenty shots, and how to review and revise output without starting from zero each time. It is tool-agnostic on purpose: the same habits work whether you are generating with Sora, Runway, Kling, Luma Dream Machine, Pika, Veo, or an image model like Flux or Midjourney used as a keyframe engine.

The central idea is simple. Treat every generation as a mini production with a shot list, a look book, and a continuity sheet. The text is the blueprint, not the decoration.

How text-to-video models actually read a prompt

What the model prioritises

Most diffusion and transformer-based video systems weigh the beginning of a prompt more heavily than the end, and concrete nouns and verbs more heavily than abstract adjectives. If your prompt opens with mood words like "cinematic" and "emotional", the model has to invent the subject. If it opens with a subject doing something in a place, the rest of the sentence becomes refinement rather than invention.

A useful mental order of priority:

  1. Subject and action — who or what, doing what, right now.
  2. Setting — where, what time of day, what weather, what era.
  3. Framing and camera — shot size, angle, movement, lens character.
  4. Lighting — source, direction, hardness, colour temperature.
  5. Style — grade, film stock, animation idiom, reference era.
  6. Micro detail — fabric weave, dust motes, background extras.

Micro detail is where prompts bloat and start contradicting themselves. Save it for later in the sentence, and only when it matters to the story.

The three layers of a professional prompt

Strong prompts stack three layers without mixing them up:

  • Narrative layer: the beat of the story. "She realises the door is already open."
  • Visual layer: what the audience sees that communicates the beat. "Wide shot, hallway, cold dawn light from a window on the left, her hand frozen on the handle."
  • Technical layer: how it is captured and finished. "Slow push-in, 35mm look, shallow depth of field, muted teal grade, 16:9."

When output disappoints, the fastest diagnosis is to ask which layer is failing. Wrong emotion usually means a narrative problem. Wrong look means a visual problem. Wrong motion, edges, or aspect means a technical problem. Fixing the correct layer is much cheaper than rewriting the whole prompt.

A reusable prompt template for script-driven video

A template prevents you from re-inventing structure every time and makes your prompts comparable, which in turn makes iteration faster. One that works across most text-to-video tools:

[Shot size and duration] | [Subject: age, wardrobe, expression] | [Action beat] |
[Location, time of day, weather] | [Camera move, lens, height] |
[Lighting: source, direction, quality] | [Style and grade] |
[Negative constraints]

Filled in, it looks like this:

Medium close-up, 6 seconds | Woman, 30s, olive raincoat, wet hair, jaw tight |
She opens an envelope and stops breathing | Rooftop car park, dusk, light rain |
Slow handheld drift left, 50mm, chest height |
Practical sodium lights behind her, soft front fill from a phone screen |
Muted filmic grade, fine grain, natural skin texture |
No text overlays, no extra fingers, no lens flare, no crowd

Notice that the prompt reads like a shot card on a real set. That is not an accident. The closer your text resembles the language a crew would use, the fewer interpretive decisions the model has to make — and the fewer places there are for randomness to leak in.

Keep a prompt sheet, not a prompt pile

The biggest operational upgrade for any team is a spreadsheet or table of shots. Columns should include: shot ID, script beat, prompt text, model used, aspect ratio, seed, reference image, take selected, and notes on what to change next time. Without this, you will regenerate the same shot six times and forget which settings produced the best version.

Turning a screenplay into shot-level prompts

Break the script into beats, not pages

A script page may contain five emotional beats and a dozen visual moments. Models work best on one beat per generation. Rewrite each scene as a numbered list of single actions: "he enters," "she notices," "the light flickers," "he leaves." Each line becomes one shot or, for complex motion, one shot split across two generations that you join in the edit.

This rewrite is where most of the value is created. If you cannot describe the beat in one sentence, the model will not be able to render it either.

Write shot lines the model can act on

Convert each beat into a shot line with five ingredients: subject, verb, object, setting, and camera. Avoid stacked verbs joined by "and," because the model often blends them into a single ambiguous motion. "She turns and drops the glass" can become a smear. Write "she turns to camera" as one shot and "the glass falls and shatters" as the next.

Use present tense. Use one camera move. Use one light source described clearly. If a shot needs two ideas, it is two shots.

Dialogue, voiceover, and on-screen text

Most video models render mouth movement poorly and dialogue not at all. The practical answer is to plan dialogue as audio-first: write the line, record or synthesise it, then generate the shot with the character either speaking off-camera, partially obscured, or filmed from behind. For voiceover, generate silent coverage — landscapes, hands, objects — and lay the narration over it. Prompt for the visual tone the narration needs rather than attempting to sync lip movement.

For on-screen text, add it in the edit. Text baked into a generation is usually warped, misspelled, or inconsistent between takes.

Keeping characters, props, and locations consistent

Build a character bible with reference frames

Consistency starts before video generation. Create a still image of each principal character — front, three-quarter, and profile — in a neutral pose with flat lighting. Write a short paragraph describing them that you paste into every prompt verbatim: age range, hair, build, wardrobe, defining details. Do not paraphrase it between shots. Small wording changes are enough to make a model rethink the face.

Where the tool supports image-to-video or reference images, feed the character still in with the shot prompt. Where it supports seeds, reuse the seed that produced the best likeness and vary only the action and camera.

Style locks for locations

Locations need the same discipline. Write one canonical description per location — architecture, palette, time of day, key light direction, recurring props — and reuse it. If a hallway has a red door in shot one, the description must say red door in shot seven.

Negative prompts: subtract rather than add

When a result has an unwanted element, resist the urge to describe its opposite. Instead, name the thing you do not want in a negative field or a final "exclude" clause: extra limbs, warped hands, text overlays, watermark-style artefacts, lens flare, crowds, modern clothing in a period scene, plastic skin, over-sharpening. Keep negative lists short and specific. A list of thirty exclusions dilutes all of them.

Plan continuity around cuts

Because each shot is generated independently, continuity problems appear at the cut, not inside the take. The cheap fix is editorial: end a shot on movement, start the next on movement, and cut on action. Motion hides small differences in costume, lighting, and framing far better than a static cut between two static shots.

Camera, motion, and pacing control

Camera vocabulary that tends to work

Models respond more reliably to established film language than to invented phrasing. Useful terms include: locked-off static, slow push-in, dolly out, tracking shot, lateral truck, crane up, handheld drift, whip pan, over-the-shoulder, low angle, high angle, drone establishing shot, macro insert. Pair each with a lens suggestion — 24mm for environment, 35mm for scenes, 50mm for intimacy, 85mm for isolation — and a height, such as eye level or chest height.

One move per shot. If you want a shot that begins wide and ends close, generate two shots and cut between them; the model will otherwise produce a drifting camera that feels unmotivated.

Prompting rhythm and pacing

Pacing is a function of shot length and cut frequency, not of the prompt alone. Decide early whether the piece is a fast-cut social edit with two-second shots or a slow brand film with six-to-eight-second holds. Then generate to that length. Asking a model for a ten-second shot when your edit needs two is wasted computation and wasted review time.

Control speed inside the frame with adverbs used sparingly: "slowly," "deliberately," "abruptly." If motion still looks wrong, reduce what happens in the shot rather than adding more descriptive language.

A production workflow from first draft to final cut

  1. Script pass. Write or adapt the script, then mark the emotional beats per scene.
  2. Beat breakdown. Convert beats into single-action shot lines with subject, verb, setting, and camera.
  3. Look development. Generate still keyframes for characters, locations, and the overall grade before touching video. Stills are faster and cheaper to iterate than motion.
  4. Shot sheet. Move everything into one table: shot ID, prompt, model, aspect ratio, seed, references.
  5. Batch generation. Generate two to four takes per shot. Review on a timeline, not in a grid — context changes what reads as usable.
  6. Assembly. Cut a rough assembly before polishing any single shot. Shots that felt weak in isolation often work in sequence.
  7. Targeted regeneration. Replace only the shots that fail in context, keeping the prompt sheet up to date with what worked.
  8. Sound and finish. Add voice, music, effects, and any on-screen text. Grade the assembled piece as a whole so the AI-generated shots and any live footage share one look.

Keeping the review at the assembly stage is the discipline that separates a two-day project from a two-week spiral.

Common mistakes and how to fix them

  • Overloaded prompts. More than about sixty words usually reduces fidelity. Split the shot.
  • Contradictory style words. "Photorealistic anime" produces mush. Pick an idiom.
  • Ignoring aspect ratio. Vertical social edits framed in 16:9 waste half the frame. Set the ratio before generating.
  • No continuity plan. Decide the character and location wording before shot one, and never improvise it mid-project.
  • Treating take one as final. Budget for multiple takes; selection is part of the craft.
  • Forgetting audio. Silent generation plus a planned soundtrack beats praying for synced dialogue.
  • Vague verbs. "She reacts" gives the model nothing. "She flinches and steps back" gives it a performance.

Choosing a model for each shot type

Different systems have different strengths, and mixing them is normal practice. Broad decision criteria:

  • Photoreal human close-ups: models with strong facial consistency and reference image support.
  • Stylised or animated sequences: image-first pipelines where you generate keyframes in an image model, then animate them.
  • Environment establishing shots: wide-capable models that hold architectural detail without warping.
  • Precise motion control: tools offering camera-motion parameters or motion brushes rather than text-only control.
  • Long continuous takes: systems with extended clip length, accepting some fidelity trade-off.

Standardise on one or two tools per project where possible. Every additional model adds a colour, grain, and motion signature you will have to reconcile in the edit.

FAQ

How long should a prompt be?
Long enough to remove ambiguity, short enough to stay readable. For most shots, two to four sentences with clear structure beat a paragraph of adjectives.

Do I need to write prompts differently for each tool?
The structure stays the same; the vocabulary shifts. Learn each tool's preferred camera terms and its handling of negatives, and keep the template identical.

Can I make a character look the same in every shot?
Yes, with three habits: a fixed written description reused verbatim, a reference still image, and a consistent seed where the tool supports it.

What about dialogue and lip sync?
Plan dialogue as audio-first. Generate silent coverage and cut away from mouths. Dedicated lip-sync tools exist for short lines, but they are best used sparingly.

How many takes should I generate per shot?
Two to four is a practical range. If none works, the prompt is wrong, not the luck.

Should I storyboard before prompting?
Yes. Even a rough sketch or a set of stills will tighten the prompts and reduce wasted generations.

Key takeaways

Write prompts like shot cards. One action, one camera move, one light idea per generation. Keep a character and location bible so continuity survives the cut. Build a shot sheet so iteration is systematic rather than emotional. Review at assembly, then regenerate only what fails in context. Do these four things and the quality of your AI video becomes a function of planning rather than luck — which is exactly how conventional production has always worked.

Alexander

Alexander