Start Free Now
Limited Time Offer: Get 50% OFF Starter & Basic Yearly Plans ๐ŸŽ‰

Text and Image to Video: A Practical AI Workflow Guide

Oct 4, 2026

Why text and image to video changed the production math

A decade ago a 30-second brand film meant a crew, a location, a lighting package, and a week of post-production. Today a two-person team with a beat sheet and a folder of reference images can produce something that holds up on a phone screen in an afternoon. The shift is not the work of one magic model. It comes from a pipeline where language and still images both act as direct inputs to motion, and where iteration costs minutes instead of days.

Text-to-video and image-to-video answer different questions. Text-to-video asks what should happen and invents the look along the way. Image-to-video asks what the frame looks like and then decides how it moves. The first is fast and exploratory. The second is controllable and repeatable. Most professional work ends up somewhere in between: a generated or photographed keyframe carries the identity of the scene, while a written prompt carries the action.

Understanding that split is the difference between a folder of pretty accidents and a finished deliverable. The workflow below is reusable across genres: plan the deliverable first, write shootable prompts, lock character consistency with reference images, pick a model per shot, direct the motion, assemble the sound, and quality-check before delivery.

Start with the deliverable, not the model

The most common failure in AI video production is starting with a tool. You open a model, type something interesting, get a beautiful five-second clip, and only then discover it does not fit the aspect ratio, the runtime, or the story you actually needed. Reverse the order. Decide what you are shipping before you generate a single frame.

The six questions to answer first

  • What is the final aspect ratio and runtime? Vertical 9:16 changes framing decisions more than any model choice.
  • Where will it be watched, and with sound on or off? Silent autoplay feeds demand on-screen text and strong visual rhythm.
  • Does it need spoken dialogue or lip sync? If yes, the voice track drives the edit, not the other way around.
  • Do characters or locations appear in more than one shot? Recurring elements mean you need reference sets and a continuity sheet.
  • Are there brand, legal, or factual constraints? Claims, logos, and likenesses must be cleared before generation, not after.
  • How many revision rounds can you afford? Agree on that number with the client, then batch feedback instead of trickling it.

Each answer removes a category of wasted generation. A vertical silent ad with no recurring characters is a completely different production from a 90-second narrative short with three speaking roles, even if both are made with the same handful of models.

Plan the shot economy

Beginners generate forty clips for a thirty-second edit. Experienced teams generate twelve and choose from those. The shortcut is a simple shot list that maps narrative beats to concrete shots with target durations.

Beat Shot Target length Source
Hook Extreme close-up of hands opening a box 2s Image-to-video from product still
Problem Wide shot of cluttered desk, handheld 3s Text-to-video
Reveal Slow push-in on the product, locked off 4s Image-to-video from render
Benefit Medium shot, person using it naturally 3s Image-to-video with reference set
Proof Macro detail, rack focus 2s Text-to-video with tight prompt
Call to action Static graphic plate, product centered 3s Generated still with motion added in the editor

With an average shot length of two to four seconds, a thirty-second spot needs roughly eight to twelve usable shots. Ten usable shots usually means generating fifteen to twenty candidates. Knowing that number in advance keeps you from over-generating in act one and rushing act three.

Budget time, attention, and compute

In practice, time splits roughly into prompt iteration, waiting for renders, and editing. The waiting is the silent killer. Two habits help: batch your generations so several clips render while you write the next shot, and keep a selects bin so you never re-watch a full session to find the good take. Generate at the highest native resolution your model supports and downscale at the end. Upscaling a low-resolution source rarely recovers detail that was never there.

Writing a script an AI model can actually shoot

A human actor can improvise around a vague line. A generative video model cannot. The screenplay for AI video sits somewhere between a script and a shot description, and its most useful order is subject, action, environment, camera, light, style.

Turn beats into shots

An emotional script line such as the character realizes she is being followed is not shootable. Two shots are. Shot one: a medium shot from behind as she walks faster, coat collar up, streetlights sliding across wet pavement, slow handheld follow. Shot two: a tight close-up on her eyes darting to a shop window reflection, static camera, shallow depth of field. Same beat, two generations, and an editor can cut between them.

A prompt template that survives editing

Use a fixed skeleton so your prompts stay comparable:

[shot size + lens] + [subject with stable descriptors] + [single action, present tense] + [environment] + [camera movement] + [lighting] + [style and grade] + [duration]

A worked example: medium-close shot, 50mm, a woman in her late thirties with close-cropped dark hair and a weathered olive jacket lifts a ceramic cup to her lips, steam rising, small cafe interior with condensation on the window, slow push-in, soft grey window light from the left, muted natural grade with slight grain, five seconds.

A weak version of the same prompt reads: cinematic woman drinking coffee, moody, epic, beautiful, 4k. It contains no action the model can time, no lens, no light direction, and no stable identifiers, which is why the face changes between takes.

Practical rules that save hours:

  • One action per clip. Two simultaneous actions produce morphing.
  • Present tense verbs beat nouns. Lifts, turns, steps, opens.
  • Avoid on-screen text in generation. Almost every model mangles lettering. Add titles in post.
  • Keep the style string identical across a project. Changing grade mid-project destroys continuity.
  • Ban vague adjectives from your own vocabulary. Cinematic, epic, and beautiful are noise.

Keep the vocabulary concrete

Build a personal phrase library of light and lens language that works: overcast daylight, hard noon sun with deep shadows, tungsten practicals, neon spill on wet asphalt, 24mm wide, 85mm portrait compression, anamorphic flare. The narrower the vocabulary, the faster you can describe a new scene, and the more consistent your results become across projects.

Building a reference image set for character consistency

Consistency is the hardest problem in multi-shot AI video, and the fix is unglamorous: more references, cleaner references, and a descriptive string that never changes.

The three-image minimum

For any recurring character, prepare at least a front view, a three-quarter view, and a profile. Add a back view if the character turns away from camera. Generate the views with the same lighting and the same wardrobe, because differences in light direction read as different people to a model. Once the set is approved, treat it as frozen. Replacing a reference mid-project is how characters suddenly change age between shots.

Lock wardrobe, hair, and age descriptors

Write a character card of six to ten fixed descriptors and paste it verbatim into every prompt that features that person: late thirties, close-cropped dark hair, weathered olive jacket, thin silver chain, calm low voice, tall and slightly stooped. The card is the single source of truth. If a jacket changes, it changes everywhere, and you regenerate every affected shot.

Location plates and lighting continuity

Generate one wide master of each location and keep it as a plate. Reference it for every setup shot in that scene so windows, furniture, and architecture stay put. Note time of day and light direction in a continuity sheet alongside the character cards. A scene that starts at golden hour should not end at noon.

Folder hygiene and versioning

Name files predictably: project, scene, character, angle, version. Keep a selects folder containing only approved references, and never let a model look at the reject folder. Drift usually starts with a good take sitting next to a bad one.

Choosing the right model for each shot

Three families of tools

Text-to-video models generate motion from language alone. They are best for establishing shots, abstract sequences, and anything without a recurring human face. Image-to-video models animate a provided frame, which makes them the workhorse for character shots and product shots because the identity of the frame is already correct. Video-to-video and motion-transfer tools take existing footage plus a style or motion reference, and they shine for dance, sport, and stylized reinterpretation.

Match strengths to shot types

Shot type Best starting point Notes
Human face close-up Image-to-video from a strong still Small motions only; big motion warps features
Product macro Image-to-video from a render or photo Lock the camera, move the light instead
Fast action Video-to-video with a motion reference Text prompts rarely produce believable speed
Stylised animation Text-to-video, then image-to-video for continuity Keep one style string for the whole piece
Establishing landscape Text-to-video Cheapest shots to iterate, generate several
Dialogue Image-to-video plus a dedicated lip-sync pass Lock the plate before syncing

Duration, resolution, and aspect ratio traps

Most models generate in five or ten second blocks. Plan edits around that rhythm, and generate slightly longer than you need so the editor has handles. Match the native aspect ratio of the model instead of cropping afterward, because cropping throws away framing decisions the prompt already made. Avoid chaining upscales: one clean upscale at the end of the edit, nothing more.

Directing the model: camera, motion, and pacing

Camera vocabulary that actually works

Models respond reliably to a small set of movements: dolly in, dolly out, truck left or right, orbit clockwise, crane up, handheld follow, rack focus, static locked-off. Pair exactly one movement with exactly one subject action. Two movements in a five-second clip produce a smeared, unreadable shot.

Control motion intensity deliberately

Every tool exposes some form of motion strength, and the temptation is always to push it high. High motion is where warping, limb duplication, and background melt appear. If a shot genuinely needs large movement, split it into two smaller moves and cut between them. Slow camera movement with fast subject movement usually looks better than the reverse.

Seeds, negative prompts, and iteration discipline

Once framing is correct, fix the seed and change one variable at a time: motion strength, then light, then phrasing. Change three things at once and you learn nothing. Negative prompts are useful for flicker, morphing hands, duplicated limbs, and on-screen text artifacts. Keep a project-level negative list instead of retyping it per shot.

Sound, voice, and dialogue that sell the edit

Voice first, picture second

Record or generate the voiceover before you lock the picture. The pace of the read dictates how long each shot can breathe, and cutting visuals first guarantees you will re-edit when the audio arrives. For synthesized voices, write for the ear: short sentences, hard consonants, no tongue twisters, and explicit pauses where you want the visuals to land.

A workable lip-sync workflow

Generate the visual without dialogue, confirm the framing and head position, then run the lip-sync pass on the locked plate. Keep the head relatively still during lines, avoid fast turns, and avoid hands near the mouth. If a line must land in a moving shot, cut to a closer, static angle for the spoken words.

Music, ambience, and effects

Build three layers. A music bed sets emotion, ambience makes the space feel real, and spot effects sell physical events: a latch click, a cup set down, a whoosh on a transition. Ambience is the layer most beginners skip, and it is the reason AI video often feels haunted and airless. Check licensing carefully on any generated track before commercial delivery.

Mix for the phone speaker

Most viewers will hear your video through a tiny mono speaker. Check the mix on a phone at low volume. Duck music noticeably under voice, keep the voice centred, and make sure no single effect is louder than the narration. If dialogue disappears at quiet playback, the mix is wrong regardless of how good it sounds on headphones.

Editing, upscaling, and quality control

The assembly pass

Cut on action and keep average shot length short: two to four seconds for social, five to eight for explainers. Order the shots from the shot list, then trim each clip to its strongest moment. In AI footage the first and last few frames are often unstable, so trimming the head and tail improves perceived quality more than any upscale.

Upscale and interpolate sparingly

Upscale once, after the cut is locked, so you are not burning time on shots you will discard. Frame interpolation can smooth motion but also produces ghosting around hands and faces; use it only when the source frame rate genuinely looks choppy on the target platform.

Quality control checklist

Before delivery, verify all of these:

  • Character descriptors identical across every prompt used in the piece.
  • Wardrobe, hair, and props unchanged between shots of the same scene.
  • Light direction consistent within a scene.
  • No warped hands, duplicated limbs, or sliding feet in action shots.
  • No garbled generated lettering; all text added in post.
  • Aspect ratio and safe margins correct for the target platform.
  • Audio peaks under control, voice intelligible on a phone speaker.
  • Luminance consistent shot to shot, no flicker on cuts.
  • Runtime matches the brief, including any end card.
  • Exported file names and versions match the client convention.

Workflow recipes by use case

Fifteen-second social ad

Six to nine shots, all vertical, written for silent viewing. Hook in the first second with a macro detail. Use image-to-video from product stills for every product appearance so packaging stays identical. Add burned-in captions in the editor. Realistic turnaround for a small team: half a day of generation and two hours of editing.

Sixty-second product explainer

Twelve to eighteen shots, horizontal, with voiceover driving the pace. Reserve text-to-video for abstract backgrounds and image-to-video for anything involving hands or the product. Build one location plate, one character card, and one style string for the whole piece. Expect one full day including sound and two client revision rounds.

Ninety-second narrative short

The hardest format, because continuity errors accumulate. Lock a reference set for every character, keep a continuity sheet per scene, and generate dialogue shots in static, close framings for easier lip sync. Budget two to three days, most of it in the edit.

Faceless channel episode

Reuse a visual template: the same three camera moves, the same grade, the same lower-third style. Consistency is the product here, not novelty. Generate in batches of ten, keep a rolling selects folder, and script for a voiceover that never needs lip sync.

Common mistakes, troubleshooting, and FAQ

Problems and fixes

Problem Likely cause Fix
Character changes between shots Drifting descriptors or no reference set Freeze a character card and regenerate affected shots
Faces wobble in close-up Too much motion for a five-second clip Reduce motion strength, split the action
Hands morph or duplicate Complex simultaneous actions Simplify to one gesture, add negatives, crop the frame
Background shifts No location plate or changed seed Reuse the plate image and fix the seed
Garbled on-screen text Models cannot render type reliably Remove text from prompts, add titles in the editor
Flat, lifeless look Vague adjectives instead of light language Specify light direction, lens, and grade explicitly
Ghosting after interpolation Too many synthesized frames Interpolate less or skip it entirely
Audio feels detached Missing ambience layer Add room tone under every scene

Frequently asked questions

How many reference images does a character actually need? Three is the practical minimum for a face that appears in multiple shots. Add views for any angle the story requires, and add a full-body plate if wardrobe matters. More variety in angles helps; more variety in lighting hurts.

Should I generate images first and animate them, or write prompts directly into video? Default to image first whenever identity matters, whether that identity is a character, a product, or a location. Use text-to-video for shots where nothing needs to be recognizable and speed matters most.

Why does my footage look artificial even when the prompt is detailed? Usually because the shot has no camera movement or the movement is too fast, and because the audio has no ambience. Add one slow camera move, add room tone, and trim the unstable first and last frames.

How long should an AI-generated shot be? Two to four seconds for social, four to six for explainer content, and up to eight for contemplative narrative moments. Longer shots expose model artifacts and lose the viewer.

Can I use generated voices and music in commercial work? That depends on the terms of the specific tool and the jurisdiction you operate in. Keep records of every generator used, and prefer tools that grant clear commercial rights at your subscription tier.

What is the fastest way to improve quality without spending more time? Trim the head and tail of every clip, unify the grade across the timeline, and add an ambience layer. Those three changes lift perceived production value more than any new model.

A short closing rule

Treat language as the director and images as the cast. When a shot fails, ask which of the two was vague. Nine times out of ten the prompt described an emotion instead of an action, or the reference set was too thin to hold an identity. Fix that, keep the pipeline boring and repeatable, and AI video stops being a novelty and starts being a production method.

Alexander

Alexander