Vente à Durée Limitée : Profitez de 30% DE RÉDUCTION sur la Création Vidéo IA de Nouvelle Génération 🎉

AI Short Film Workflow: From Text and Images to Video

Sep 14, 2026

Treat AI Video Like a Pipeline, Not a Slot Machine

Most people's first encounter with AI video generation looks like this: they type a poetic paragraph into a text-to-video box, wait ninety seconds, and get back a beautiful four-second clip of something vaguely related to what they imagined. It looks impressive in isolation. Then they try to build a sixty-second story out of it and everything collapses. Faces change between shots. Camera angles contradict each other. Lighting jumps from golden hour to fluorescent office and back again.

The problem is rarely the model. It is the missing production layer. A short film, even a ninety-second one, has structure: a premise, a point of view, a rhythm of wide shots and close-ups, a consistent world. Generative models produce fragments. Filmmaking is the act of assembling fragments into meaning.

So the practical approach is to borrow the discipline of traditional production and apply it to a faster, cheaper, stranger toolbox. Lock a script. Build a shot list. Establish a visual bible. Generate stills before motion. Treat animation as the final step rather than the first. This guide walks through that pipeline in the order you would actually do the work, with the decision points where most projects go wrong.

Choosing Your Generation Path

Before writing a single prompt, decide which of three paths your project will take. This decision shapes everything downstream: your budget of time, your tolerance for retries, and how much control you keep.

Text to video

You describe a shot in words and the model produces motion directly. This is the fastest path and the least controllable. It works well for abstract sequences, landscapes, atmosphere, and establishing shots where no specific character needs to remain recognizable. It works poorly when you need the same face in six different shots.

Use text to video when: the subject is environmental or generic, you need speed over precision, or you are storyboarding and just want to test a mood.

Image to video

You generate or supply a still frame, then ask the model to animate it. Because the composition, wardrobe, and lighting are already fixed in the still, the motion model has far less room to invent something wrong. This is the backbone of most serious AI short films.

Use image to video when: character consistency matters, the shot requires a specific composition, or you need to control framing precisely.

Hybrid

Most finished projects are hybrid. You might use text to video for a drone-style opening, image to video for every dialogue scene, and simple still-image compositing with camera moves for moments where motion would hurt more than help.

A useful rule: if a shot needs to match another shot, generate a still. If a shot stands alone, text to video is fine.

Stage One: Write the Script for Your Constraints

AI short films fail most often at the script stage, and not because the writing is bad. They fail because the script describes things no generative pipeline can hold together.

Write with these constraints in mind.

Limit locations. Three locations in ninety seconds is ambitious. Two is comfortable. One is a superpower, because a single location lets you reuse lighting references and background stills across many shots, which instantly improves visual coherence.

Limit speaking characters. Every additional character is another identity you must maintain across every shot they appear in. Two to three is a realistic ceiling for a first project.

Prefer visual action over dialogue. Voice is easy to generate but hard to sync convincingly to lips. Write scenes where meaning comes from gestures, objects, and cuts. A character picking up a photograph and turning away carries more weight than a monologue.

Write shots, not prose. Instead of "Maya wanders through the market, lost in thought," write the shot list inside the script: wide of Maya entering the market, medium of her hand brushing fabric, close on her eyes drifting to a stall, insert of a wooden carving. Each line becomes one generation.

Keep a clock. A 90-second film at 2 to 4 seconds per generated shot is roughly 25 to 45 shots. That is a real number you can plan against. If your script implies 90 shots, cut it in half before you start.

Stage Two: Build a Visual Bible and Shot List

The visual bible is a single document that answers every aesthetic question before you spend time generating. It typically contains:

  • Palette: three to five named colors, described in plain language ("cold slate blue shadows, warm amber practicals, desaturated mid-tones").
  • Lighting logic: the direction and quality of light for each location. "Interior diner: overhead fluorescent, slightly green, hard shadows."
  • Lens and framing language: do you shoot wide and anamorphic, or tight and handheld? Pick two or three shot types and reuse them.
  • Character sheets: one still per character, front-facing, neutral background, plus a written description of hair, wardrobe, age, and build that you paste into every prompt.
  • Texture references: grain, film stock feel, and level of realism (photoreal, painterly, animation).

Then build the shot list as a table with columns for shot number, description, shot size, location, characters present, duration, and status. This table becomes your production control panel. When a shot fails repeatedly, mark it and move on rather than stalling the whole film.

A practical tip: number shots in tens by scene. Scene one gets 10, 20, 30; scene two gets 40, 50, 60. You can insert new shots later without renumbering everything.

Stage Three: Generate Stills and Hold Character Consistency

This is where the film is actually made. Everything after this is execution.

Generate a character reference first

Create one strong, clear image of each character before generating any scene. Front-facing, neutral background, even light. Then examine it critically: is the wardrobe describable in one sentence? If not, simplify it. A character wearing a plain gray coat is reproducible. A character wearing a coat with six visible details is not.

Reuse the same description block verbatim

Copy and paste the exact character description into every prompt that includes them. Do not paraphrase. Small wording changes produce visible identity drift. Keep a text file with description blocks and paste from it.

Use the still as a reference image

Most image-to-video and image generation tools accept a reference image alongside a text prompt. Feed the character sheet in as a reference and describe only what changes: pose, framing, action, environment. Let the reference handle identity.

Solve lighting before you solve faces

Oddly, inconsistent lighting breaks the illusion of a coherent film faster than inconsistent facial detail. If your scene is lit by a single window, every shot in that scene needs window light from the same direction. Write that into the shot list and into every prompt.

Generate more than you need

For every shot, generate three to six variants. Choose the best. Keep the second-best as a backup in case the animation stage fails on your first pick. Deleting results is cheap; re-generating from scratch after you have moved on is expensive in time and momentum.

Stage Four: Animate the Stills

Now you turn stills into motion. The core rule is that the less you ask the model to change, the better the result.

Write motion prompts, not scene prompts

The still already describes the scene. Your animation prompt should describe only movement: "slow push in, subject turns head slightly toward camera, steam rising from cup, subtle fabric movement." If you describe the whole scene again, the model may re-render it and lose the identity you carefully built.

Match motion to shot size

  • Wide shots: slow drift, crane move, or gentle parallax. Avoid fast motion; wide shots with rapid movement expose model artifacts.
  • Medium shots: small pushes, slight handheld sway, subject walking toward camera.
  • Close-ups: micro-motion only. A blink, a breath, a slight tilt of the head. Aggressive motion in close-up is where AI video looks most artificial.

Keep clips short

Two to four seconds is the sweet spot for most tools. Longer clips drift, morph, and lose coherence. You will assemble them in the edit anyway, so short is not a limitation; it is a workflow.

Interpolate for smoothness, not for length

Frame interpolation and motion smoothing can rescue a slightly choppy clip. Use them sparingly. Over-interpolated footage develops a soap-opera quality that reads as artificial.

Know when to stop retrying

Set a retry limit per shot, typically three attempts. If it still fails, the shot is probably wrong for the medium. Replace it with a different framing, a static image with a slow zoom, or simply cut it. The audience will never miss a shot that was never in the film.

Stage Five: Sound, Voice, and Music

Sound is the single highest-leverage upgrade for AI short films, because audiences forgive imperfect visuals far more readily than they forgive bad audio.

Voice

Generate dialogue line by line, then edit the timing. Do not try to match voice to an existing animated mouth; instead, cut to reaction shots, over-the-shoulder angles, or objects while the line plays. This is standard documentary technique and it solves a technical problem elegantly.

Ambience

Lay a continuous ambient bed under each scene: room tone, distant traffic, wind, crowd murmur. This single track does more to make cuts feel intentional than any transition effect.

Foley

Add specific sounds for specific actions: a cup touching a table, a door latch, footsteps on gravel. Even rough Foley sells the reality of a shot.

Music

Choose a single mood and stay with it, or use two contrasting cues tied to the story's turn. Avoid constant music; silence before a key moment is more powerful and cheaper to produce.

Mix

Balance dialogue, ambience, and music with dialogue always intelligible. Normalize loudness across the film so no scene jumps out. A simple three-track mix in any editor is enough.

Stage Six: Edit, Color, and Finish

AI-generated footage rarely matches perfectly straight out of the tool, so the edit does the unifying work.

Assemble rough first. Cut on action and on movement. If a character raises a hand in shot A and the next shot begins with a hand already raised, the cut feels natural even if the two shots were generated weeks apart.

Trim aggressively. Cut the first and last half-second of most generated clips; that is where morphing and drift concentrate.

Unify color. Apply a single look across the whole film: slight desaturation, a shared color temperature, a gentle contrast curve. This one step does more for coherence than regenerating footage.

Add grain and texture. A subtle grain layer hides minor artifact differences between clips and gives the film a consistent surface.

Check pacing against the script. Read the film as a silent cut before adding score. If it holds attention without music, the edit works.

Export and watch on a phone. Most short-form audiences watch vertically on small screens, where details vanish and contrast flattens. Watch your cut there before declaring it finished.

Mistakes That Sink AI Short Films

Starting with animation. Generating motion before you have locked stills means every identity problem appears multiplied. Stills first, always.

Overloading prompts. Five subjects, three actions, and elaborate lighting in one prompt produces mush. One idea per shot.

Ignoring continuity of light. Audiences track light direction subconsciously. Mismatched light reads as a broken film even when faces match.

Chasing perfection on one shot. Ten attempts on shot 12 while shots 13 through 40 remain ungenerated is how projects die. Move on, then revisit.

No rough cut until the end. Assemble a rough cut as soon as you have a third of your shots. Problems that look like generation failures are often editing problems.

Neglecting audio. Beautiful footage with thin, silent audio feels like a demo. Rough audio under strong visuals feels like a film.

Skipping the script. The fastest way to waste a week is to generate without knowing what the film is about.

FAQ

How long should a first AI short film be?
Sixty to ninety seconds. That is roughly 20 to 40 shots, which is enough to learn the entire pipeline without exhausting your patience.

Do I need image generation skills to do this?
You need the ability to describe an image clearly and to evaluate results critically. Technical skill matters less than consistency: using the same description blocks, the same palette, and the same lighting logic across every asset.

What if my character keeps changing between shots?
Build a dedicated character reference still, reuse the identical description text verbatim, and pass the reference image into every generation. Also check your lighting: identity drift is often actually lighting drift.

Should I animate everything?
No. Static shots with slow camera moves are legitimate filmmaking and often more effective than generated motion. Use animation where movement carries meaning.

How do I handle dialogue scenes?
Generate voice separately, then cover the lines with reaction shots, inserts, and over-the-shoulder angles rather than attempting lip sync. It is faster, more reliable, and looks more professional.

Where does color correction fit in?
At the end, applied globally. One consistent look across every clip is the fastest way to make disparate generations feel like one film.

How do I keep a project from stalling?
Set retry limits, keep a shot status table, and assemble a rough cut early. Momentum is a production asset; protect it with rules rather than willpower.

The pipeline is not glamorous, but it is the difference between a folder of impressive clips and a short film someone actually watches to the end.

Alexander

Alexander