Why text-to-video changes the production math
For most of film history, the cost of a shot was dominated by logistics: crew, location, lighting, permits, travel, and the hours burned resetting between takes. AI video generation does not make those costs disappear, it relocates them. In a generated workflow, the expensive resource is not equipment. It is iteration time and judgment. A single creator with a clear shot list can now produce a three-minute narrative short that would previously have required a small crew and a rented location.
That shift has three practical consequences worth understanding before you open any tool.
- Pre-production matters more, not less. When a model can produce twenty variants of a shot in minutes, the bottleneck becomes knowing which variant is correct. That requires a locked script, a clear visual reference, and a shot list.
- Continuity becomes the main technical problem. Anyone can generate one beautiful four-second clip. Making twelve clips feel like the same film, in the same place, with the same person, is the actual craft.
- Sound carries more weight than image. Viewers forgive slightly soft visuals. They do not forgive dialogue that misses sync, ambience that cuts abruptly, or music that fights the scene.
Modern models have improved markedly in prompt adherence, physical plausibility, camera language, and native audio. The ceiling of what a solo creator can make has risen accordingly. It also means the gap between amateur and professional output is now almost entirely a function of workflow discipline rather than access to tools.
The end-to-end AI video workflow at a glance
Treat your project as seven stages, not one continuous generate button.
- Format decision. Decide duration, aspect ratio, and platform before writing anything. A 60-second vertical piece and a five-minute horizontal short are different crafts with different shot economies.
- Script and beat sheet. Write the story in beats of visible action, roughly one beat per four to eight seconds of screen time.
- Shot list and lookbook. Convert beats into shots, then collect six to twelve reference images that define palette, lensing, and texture.
- Look development. Generate three or four test frames or short clips to lock a visual direction before committing to the full shot list.
- Generation. Produce clips in batches, keep a log, and change one variable at a time.
- Assembly. Edit a rough cut with placeholder audio before you perfect any single shot.
- Sound and finishing. Replace placeholders, mix, colour-match, and export.
The most important thing about this list is that it is a loop, not a line. You will return to generation during assembly, usually for pickup shots, inserts, and any moment where the cut exposes a continuity gap. Budget for it: assume 20 to 30 percent of your generation time happens after the first rough cut.
Scripting for generation: prompts a model can actually shoot
Video models do not understand interior states. They understand visible behaviour, physical space, and camera movement. The single biggest upgrade you can make to your prompting is translating emotion into observable action.
Converting an emotional beat into observable action
"She realises she has been betrayed" is not filmable in one shot. "She stops mid-step, looks down at her phone, her jaw tightens, then turns and walks back the way she came" is. The second version gives the model a beginning, a change, and an end, which is what a shot is.
Write your beats as a sequence of physical events, then decide which events deserve their own shot. If a beat needs three sentences of prose to explain, it probably needs three shots rather than one clever prompt.
Prompt anatomy: subject, action, camera, light, lens, style
A reliable prompt structure runs in this order:
- Subject: who or what, with one or two identifying details that will repeat across shots, for example "a woman in a charcoal wool coat, dark bob, small scar above the left eyebrow."
- Action: the physical event of the shot, written in present tense.
- Camera: shot size, angle, and movement, such as "medium shot, slow push-in" or "wide static shot, slightly low angle."
- Light and time: "overcast late afternoon," "single practical lamp, warm," "hard noon sun."
- Lens and texture: "35mm, shallow depth of field," "handheld, slight grain."
- Style: keep it short and identical across the whole film. Two or three words are usually enough.
A worked example: "Medium shot, a woman in a charcoal wool coat with a dark bob stands at a rain-slicked bus stop. She glances left, exhales, pulls her collar up. Slow push-in, overcast late afternoon, cool palette, 35mm, shallow depth of field, restrained realist style."
Every element is doing a job. The subject block is reusable across the film. The action block is unique to this shot. The camera, light, lens, and style blocks are copied verbatim into every prompt, which is what makes the finished film look like it was shot by one person on one day.
Duration budgeting and shot economy
Most clips land comfortably between four and eight seconds. Plan your film around that rhythm. A 90-second piece is roughly 18 to 25 shots, which is a substantial amount of generation work, so design shots that carry their own micro-arc rather than writing forty shots of two seconds each.
Two rules save enormous time. First, avoid shots that require precise hand interaction or readable text, since these remain the weakest areas of most models. Second, avoid long dialogue shots unless your tooling handles lip sync well. A shot of a character listening is easier to generate, cheaper to iterate on, and often better cinema than a talking head.
Pre-production: shot lists, lookbooks, and continuity
Building a shot list that survives generation
A shot list is a contract with yourself. Keep these columns: shot ID, beat, description, duration, shot size and camera move, location, characters present, audio notes, and status. The status column (planned, generating, approved, replacement needed) is what keeps a twenty-shot project from turning into chaos two hours in.
Review the list for coverage balance before you generate anything. If every shot is a medium of the same character, the edit will feel flat no matter how good the individual clips are. Add wides for geography, inserts for texture, and a couple of over-the-shoulder frames for rhythm. A useful ratio for a first short is roughly 50 percent coverage of the main action, 30 percent reactions, and 20 percent environment and inserts.
Look development: reference frames and palette locks
Collect six to twelve reference images before generating anything: film stills, photography, paintings, or frames you generated yourself. From them, extract a written look sheet covering palette, contrast, grain, lens range, and lighting logic. This sheet is what you paste into every prompt.
Consistency across a film comes from repeating the same style language verbatim, not from describing the look freshly each time. If you improvise a new way to say "cool overcast light" in shot nine, shot nine will look like it belongs to a different film.
Generation discipline: batching, seeds, and iteration
Professional-looking output usually comes from volume plus selection, not from one magic prompt.
- Batch three to five variants of every shot before evaluating any of them. Judging a single clip is misleading, because you cannot tell whether a flaw is systemic or a one-off.
- Use seeds where the tool supports them. Lock a seed when the composition is right, then change only the action wording.
- Change one variable at a time. If you alter subject, camera, and lighting simultaneously, you learn nothing from the result.
- Timebox each shot. Twenty focused minutes per shot is a reasonable cap. If you are still not close, the prompt needs restructuring, not more rolls.
Keep a generation log: shot ID, prompt version, seed, tool, and verdict. Without it you will rediscover the same failures on the same shots a week later, and you will burn hours re-testing ideas you already rejected.
Reading failures like a director
Most failures fall into recognisable categories, and each has a specific remedy.
- Wrong action: the model ignored a verb. Split the action into a simpler single event, or reduce the number of subjects in frame.
- Morphing and anatomy drift: usually caused by too much simultaneous motion or an extreme camera move. Simplify the action and slow the camera.
- Camera drift: the model adds a move you did not ask for. State "static shot" explicitly and remove any word that implies movement.
- Style drift: the palette shifted between shots. This is almost always a prompt inconsistency, so check that your style line is identical to the previous shot.
- Text or logo artifacts: remove any requirement for readable text and add it in post instead.
When to re-roll, rewrite, or switch tools
Use a decision order rather than reacting to each bad clip. Re-roll when the composition and motion are right but small details are off. Rewrite the prompt when two or three consecutive variants fail in the same way. Switch tools only when a specific capability is missing: reference conditioning, native audio, longer duration, or stronger camera control. Tool hopping before rewriting is the most common waste of time in AI video production, because it resets your learning without fixing the underlying prompt problem.
Consistency: characters, wardrobe, and locations
Consistency is the craft problem that separates a demo from a film.
Wardrobe and prop language locks
Write a short character sheet for each person: age range, build, hair, one distinguishing feature, and full wardrobe description. Then copy that exact phrasing into every prompt where the character appears. Do not paraphrase. "Charcoal wool coat" and "dark grey coat" will produce two visibly different garments, and the audience will read that as a continuity error even if they cannot name it.
The same discipline applies to props. If a red mug matters in shot three and shot nineteen, describe it identically both times, including its position on the table if the framing matters.
Techniques that improve continuity
- Reference conditioning: many tools accept one or more reference images, or a first and last frame, to anchor appearance and composition. Multi-reference input is far more reliable than describing a face in words.
- Image-to-video: generate and approve a still, then animate it. This gives you control over casting and framing before motion enters the equation.
- Coverage strategy: for characters that are hard to hold, use wides, silhouettes, over-the-shoulder frames, and reaction shots where the face is partly obscured. Audiences read continuity from wardrobe and silhouette more than from facial detail.
- Location locks: describe each location once in a location card, including time of day, weather, and the light source, then reuse it exactly.
If a character must appear in close-up, generate that close-up early in the process. If it does not hold up, restructure the scene so the close-up is not required. Discovering this during assembly is far more expensive than discovering it during look development.
Audio: dialogue, ambience, and the sound design layer
Sound is where most AI video projects are won or lost, and it is usually treated as an afterthought.
Dialogue workflows
Three approaches are viable, and they suit different projects.
- Native generated audio inside the video tool. Fastest, and lip sync is handled automatically, but control over performance nuance is limited.
- Separate voice generation plus a lip sync pass. More control, better for scripted dialogue, and much easier to revise when a line changes.
- Recorded human voice. Still the highest quality option for narration and lead performances, and often faster than fighting a synthetic take into shape.
Whatever you choose, generate or record dialogue per line rather than per scene. Editing individual lines is trivial. Re-cutting an entire scene because one sentence changed is not.
Ambience and music
Build an ambience bed for every location: room tone, traffic, rain, crowd, wind. Keep the bed consistent across every shot in that location, and let it run slightly under the cut points so transitions feel seamless rather than stitched.
Select music after the rough cut, when you know the actual rhythm of the edit. Choose one track, keep it simple, and duck it under dialogue rather than switching tracks repeatedly. Practical mix targets to aim for: dialogue as the loudest element, ambience roughly 15 to 20 dB below dialogue peaks, and music below both. If you can clearly hear the ambience during dialogue, it is too loud.
Editing and finishing the cut
Cutting for continuity between mismatched clips
Generated clips rarely match perfectly, and the edit is where you hide that. Cut on motion, so the eye follows a turn, a hand movement, or a step rather than the image change. Use reaction shots as bridges between two shots that do not sit together. Insert a texture shot of hands, environment, or an object when you need a beat of separation. Trim slightly shorter than the generated clip, because the first and last frames are usually the least stable.
Colour, grain, and unification
If clips came from different tools or different prompt versions, they will differ in contrast and colour temperature. A simple finishing pass fixes most of it.
- Match black levels and white balance across the timeline before applying any creative look.
- Apply one subtle look to the whole film instead of grading shot by shot.
- Add light grain to unify texture, especially when mixing sources.
- Use a single consistent crop or letterbox treatment.
Unification, not per-shot perfection, is what makes a film feel deliberate. Export at the highest quality your platform accepts, then check the final file on a phone screen as well as a monitor, because most viewers will see it small.
Tool selection: criteria that actually matter
Rather than chasing the newest model, evaluate tools against the specific film you are making.
- Control surface: can you supply reference images, first and last frames, and explicit camera instructions? For narrative work, control beats raw realism.
- Duration and resolution: does it produce clips long enough for your shot plan without awkward stitching?
- Native audio: does it generate usable sound, and can you replace it cleanly if not?
- Consistency features: seed control, character references, and style locking.
- Iteration speed: how quickly can you produce five variants? Iteration speed matters more than peak quality, because you will generate dozens of clips per finished minute.
- Cost model: think in usable clips per hour of work rather than headline rates.
- Rights and licence terms: confirm what you can publish commercially before building a project around a tool.
A practical evaluation test: take one dialogue-free twenty-second scene and generate it in three candidate tools. Count how many variants you needed per approved shot. That number, multiplied by your shot count, is your real production cost.
Quality control, common mistakes, and frequent questions
Before publishing, run this checklist:
- Watch the film with sound off. Is the story still legible?
- Watch it with your eyes closed. Does the sound hold up on its own?
- Check every shot for anatomy, extra limbs, and morphing at the edges.
- Confirm wardrobe, props, and time of day are consistent.
- Verify dialogue sync at the start and end of every line.
- Check audio levels on headphones, laptop speakers, and a phone.
- Watch the final export, not the timeline, end to end without stopping.
Common mistakes
Writing a feature-length script. Start with 60 to 90 seconds, finish it, then scale up.
Generating before the look is locked. Test frames first. A locked look typically halves the number of wasted variants.
Perfecting one shot. A beautiful shot four that breaks the rhythm of the film is worse than an adequate shot four that cuts well.
Ignoring sound until the end. Build ambience as you go, because a missing ambience bed will change how you cut.
Editing on the timeline instead of the cut. Assemble fast, watch the whole thing, then fix what the full watch reveals.
Frequently asked questions
How long does a one-minute AI film take? With a locked script and shot list, eight to fifteen hours across generation, sound, and editing is realistic for a first attempt, and less as your prompt library grows.
Do I need to be a video editor? Basic editing literacy helps enormously. Rhythm, coverage, and sound balance are craft skills that generation tools do not replace.
Should I use image-to-video or text-to-video? Use image-to-video for anything with a recurring character or a precise composition. Use text-to-video for establishing shots, inserts, and texture.
How do I stop characters changing between shots? Lock wardrobe language verbatim, use reference images, prefer wides and silhouettes, and generate close-ups early so you know what you are working with.
Can I mix output from multiple tools in one film? Yes, with a unifying colour, contrast, and grain pass. Keep the style language consistent across tools so the underlying look does not drift.
What kills an AI video project fastest? The absence of a shot list. Without one, every clip becomes an isolated experiment and nothing assembles into a film.


