Why text-to-video belongs in a normal production pipeline
For a long time, the distance between a written concept and a finished product video was measured in weeks: storyboards, location scouting, casting, a camera crew, a colourist, a sound designer, and a round of client notes that undid half of it. Text-to-video models collapsed a large part of that pipeline into a browser tab. You describe a shot, and a model returns a few seconds of moving footage that is frequently good enough to sit inside a thirty-second explainer.
The catch is that "good enough" carries a lot of weight in that sentence. One beautiful clip is easy. A sequence of eight clips that share one character, one location, one lighting scheme, and one visual tone is genuinely difficult. That is where the craft now lives: not in generating footage, but in directing consistency.
This guide walks through a complete, model-agnostic workflow for producing a realistic product or service explainer from a text brief. It covers script design, visual bibles, prompting structure, engine selection by shot type, sound, assembly, and the quality checks that separate an advertisement from a demo reel. It is written for freelancers, in-house marketing teams, and small studios who need repeatable output rather than lucky one-offs.
It is worth being honest about fit. This approach works extremely well for lifestyle context, mood, environment, hands-and-objects beauty shots, and conceptual sequences. It works less well for engineering-accurate depictions: exact dimensions, precise material textures under controlled light, a real interface with real text, or the face of a real person. In those cases, hybrid production — shoot or render the product properly, generate the human and environmental footage around it — will always outperform a fully synthetic timeline.
What realism actually means when a model renders your product
Realism is not a single dial. It is four separate layers, and each one fails differently. When a viewer says a clip "looks AI," one specific layer has usually broken, and knowing which one saves you from re-rendering everything.
Layer one: physical realism
Physical realism is about how matter behaves. Fabric folds. Liquid finds its level. A hand grips a mug with pressure instead of hovering near it. Shadows fall away from the light source rather than toward it. Modern models handle cloth, water, and smoke well. Hands, reflective surfaces, and fine mechanisms — zips, hinges, jewellery clasps, watch bezels — remain the classic failure points. If a shot depends on a small mechanism reading correctly, budget extra takes or design around it.
Layer two: temporal realism
Temporal realism is how motion persists across frames. Objects should not teleport, breathe at inconsistent rates, or subtly morph while the camera moves. Long shots expose this; short ones hide it. If you have a choice, break an eight-second action into two four-second beats and cut on the motion. Temporal artefacts are also more visible in slow motion, so resist the urge to stretch a clip to fill time.
Layer three: continuity realism
Continuity is what makes a sequence feel like one story rather than a mood board: the same jacket, the same kitchen, the same overcast light, the same colour temperature. This is the hardest layer to control because it depends far more on your reference material than on the model's raw talent. Continuity is a pre-production problem, not a rendering problem.
Layer four: performance realism
The final layer is behavioural. Real people glance off-camera, hesitate before speaking, fidget with a sleeve, shift their weight. A synthetic performer who delivers a perfect, continuous, eye-locked monologue reads as uncanny even when every pixel is flawless. Deliberate imperfection is a production decision, not a compromise. A half-second of looking away can do more for believability than another hour of rendering.
Pre-production: write the script as a shot list, then lock a visual bible
Most people start by writing marketing copy and then try to illustrate it. That order produces generic footage, because marketing language describes feelings rather than things. Reverse it: write the shots first, then let the voiceover carry the copy.
The five elements of a production-ready shot line
A usable AI shot line contains five things, in the same order every time:
- Shot number and duration (for example, 3s)
- Subject and action (who does what, in one clause)
- Setting (location, time of day, weather)
- Camera (framing and movement)
- Light and mood (one or two adjectives, no more)
Here is the difference in practice. A weak line reads: "Show the app being useful for busy professionals." A strong line reads: "Shot 4, 3s — woman in her thirties taps a phone screen while walking through a bright co-working corridor, medium close-up, slow stabilised push-in, soft window light, calm and competent." The first cannot be rendered. The second can, and it can be rendered again tomorrow by someone else on your team.
Design for eight seconds or less
Short clips are your friend. A model asked for six seconds of continuous action tends to produce something coherent. The same model asked for fifteen seconds will often drift, simplify, or hallucinate a new background halfway through. Build your piece from many short beats and assemble them in the edit. Viewers read the result as one continuous scene if the continuity holds.
One useful constraint: give each clip a single primary action. If your prompt asks for a walk, a turn, a hand gesture, and a camera move in four seconds, something will break. Choose the one action the beat needs and let the cut supply the rest.
What belongs in a visual bible
A visual bible is a one-page document that locks the look so every prompt returns compatible footage. Without it, you generate attractive clips that cannot be edited together. With it, even mediocre clips will cut. Keep it to a single page and treat it as a contract.
- Character sheet: age range, build, hair, wardrobe with exact colours, two or three reference stills.
- Location sheet: room, wall colour, window direction, key furniture, two reference stills.
- Palette: three colour values plus one accent.
- Lighting rule: for example, "soft diffused daylight from camera left, never hard sun."
- Lens rule: for example, "35mm equivalent, shallow but not extreme depth of field."
- Grade rule: contrast level, saturation level, warmth.
Generate the still before the motion
Do not jump straight to video. Use an image model to produce a still of your character in your location, then iterate on that still until it is exactly right. Use it as a reference image for every generation in that scene. Reference-driven generation is dramatically more consistent than text-only generation, and it gives your client or stakeholder something concrete to approve before you spend time rendering motion. In practice, an extra hour on references routinely removes several hours of re-rolling.
Prompting as an assembly line, not an act of inspiration
Once your bible exists, every prompt should be assembled from the same blocks in the same order. Structure beats vocabulary, and repeatability beats clever wording.
[reference image] + [subject & wardrobe] + [action] + [location] + [camera move] + [light] + [lens/grade] + [negative constraints]
Keep this template in a text file and paste it for every shot in a scene. The only blocks that should change are the subject action and the camera move.
Camera language models actually understand
Models respond reliably to a fairly small vocabulary: slow push-in, slow pull-back, lateral tracking, handheld, static tripod, low angle, high angle, over-the-shoulder, and rack focus between foreground and background. Vague instructions such as "cinematic camera work" produce vague movement. Name the axis of motion and roughly the speed, and describe where the camera starts and where it ends.
Lighting cues that do the heavy lifting
The fastest way to make generated footage look premium is disciplined lighting language. "Soft window light with gentle falloff" reads as considered. "Dramatic lighting" reads as student film. Avoid mixing multiple light sources in one shot unless the story genuinely needs it, and keep the direction of your key light consistent across the whole sequence so the edit does not feel like a collage.
Treat negative constraints as part of the prompt
Most tools accept a negative field. Keep a reusable list: no text overlays, no logos, no extra fingers, no warped hands, no lens flare, no crowds, no heavy shadows under the eyes, no fast cuts. Copy it into every generation so you are not troubleshooting the same artefacts shot after shot. When a new artefact appears, add one line to the shared list rather than fixing it locally.
Change one variable at a time
When a shot is wrong, change exactly one thing — camera, then light, then wardrobe. Changing three variables at once produces a better-looking image with no idea why, which is useless on shot six when you need to reproduce the result. Keep a simple log: shot number, what you changed, which take you kept. Two columns in a spreadsheet is enough, and it turns a lucky find into a repeatable technique.
Choosing the right engine for each shot type
The temptation is to pick one model and use it for everything. In practice, different shots suit different engines, and a mixed pipeline usually looks better than a single-model one.
Broad capability models for human performance
Flagship general video models excel at natural human motion, facial expression, and multi-subject scenes. Use them for dialogue-adjacent beats, walking shots, and anything where a person must look like a person. Expect slower turnaround and tighter clip length limits, and plan your shot list around those limits rather than discovering them mid-render.
Image-to-video specialists for product beauty shots
When the subject is an object — a bottle, a watch, a laptop, a sneaker — image-to-video pipelines built around a hero still give you precise control. You choose the exact frame you want to animate, then add a small, believable motion: a slow rotation, a light sweep across glass, condensation forming, steam rising. Small motion with accurate materials beats large motion with smeared detail every time.
Stylised engines for transitions and abstract sequences
Some models are better at graphic, high-contrast, or stylised motion. Use them for transitions, background loops, and abstract B-roll where photorealism is not the goal. Forcing a stylised engine into photoreal product shots wastes hours and never quite lands.
Locally runnable models for volume and privacy
If you need dozens of variants of the same shot, or you cannot send client assets to a hosted service, locally runnable models are worth the setup cost. Quality is lower, but iteration speed and data control are higher. Many studios use hosted models for hero shots and local models for background plates and looping textures.
Selection criteria that actually change the outcome
| Criterion | Why it matters |
|---|---|
| Maximum clip length | Determines how finely you must cut your action into beats |
| Image reference support | The single biggest factor in character continuity |
| Motion control options | Camera presets reduce prompt guessing |
| Output resolution and aspect ratio | Vertical and square crops should be planned, not discovered |
| Commercial usage terms | Non-negotiable for client work |
| Iteration speed | Determines whether you can afford twenty takes |
| Negative prompt support | Saves you from fighting recurring artefacts |
A pragmatic approach for a thirty-second piece: shortlist two engines, test both on the same hero shot with the same reference image, and let the test decide. Ten minutes of comparison beats an hour of opinion.
Sound carries more realism than pixels
Audiences forgive soft footage far more readily than bad sound. If your budget is limited, spend it here after the visual bible.
Voiceover
Modern text-to-speech is convincing enough for most product explainers when you write for the ear: short clauses, concrete nouns, no stacked adjectives. Generate three takes at slightly different pacing, then choose per sentence rather than per paragraph, because pacing problems are local. Small pauses between sentences are easier to add in the edit than to remove, so favour a slightly slower read.
Room tone and effects
Every generated clip is silent, which is why raw footage feels hollow. Add a continuous room tone under the whole sequence — a faint hum, distant traffic, an air-conditioning hiss. Then layer specific effects: a fabric rustle on a costume change, a subtle whoosh on a transition, a soft click when a lid closes. These tiny sounds do more for believability than another hour of rendering, and they cost almost nothing in time.
Music
Choose an instrumental bed with no dominant melody and no obvious drop. If the music has a hook, it competes with your voiceover. Duck it six to ten decibels under speech and let it swell only in the two seconds around your product reveal. One track is enough for a thirty-second piece; two competing tracks signal indecision to the viewer even if they cannot name it.
Edit, grade, and finish like a normal shoot
Editing is where a sequence of clips becomes a film. Work in a standard editor and treat the generated clips exactly as you would rushes: log them, label them, cut them.
Cut on motion
Place your cuts during movement rather than at rest. A hand reaching, a turn of the head, a door opening. Motion masks small continuity errors because the eye is following the action rather than comparing frames. If a clip has a weak final half-second, cut before it rather than trying to fix it in post.
Match colour across clips
Even with a locked bible, clips will drift in temperature and contrast. Use a colour match tool, then apply one grade over the finished sequence — a soft film-style curve plus a slight warm shift is a reliable neutral. Consistency of grade is what makes a mixed-source timeline look like one shoot rather than three vendors.
Add a subtle camera layer
A very light film grain plus a barely perceptible handheld drift over the whole timeline removes the pristine, synthetic stillness that plagues generated output. Keep it subtle; heavy grain looks like a filter, not like film.
Overlay graphics sparingly and always in the editor
Text overlays, interface mockups, and lower thirds should be added at final resolution in the edit — never generated inside the model. Generated lettering is unreliable, changes shape between frames, and will need replacing anyway. The same applies to prices, legal lines, and any claim you would have to defend.
Common mistakes, and the checklist that catches them
The five mistakes that waste the most time
The first is writing prompts as marketing copy. "Premium, innovative, seamless experience" gives a model nothing to render. Translate every adjective into something visible: premium becomes dark walnut and brushed steel; seamless becomes a single unbroken camera move.
The second is generating too few takes. Professional pipelines assume a high rejection rate. Ten to twenty generations per finished shot is normal, and planning for it prevents the temptation to accept a mediocre clip because you are tired of prompting.
The third is ignoring the motion budget. Models have a limited amount of coherent movement per clip. Give each clip one primary action and let the cut supply the rest.
The fourth is over-stylising. Heavy neon, extreme lens flares, and aggressive slow motion date quickly and signal "demo" rather than "advertisement." Restrained, well-lit, ordinary scenes age better and convert better.
The fifth is skipping the edit. Treating generation as the whole job produces a slideshow. The cut, the grade, and the sound design are where the piece earns trust.
Pre-export quality checklist
Run this list before you export. It catches most of the issues that make viewers distrust a clip without being able to say why.
- Hands and fingers: count them in every shot with visible hands.
- Eyes: check for drifting pupils, uneven blinks, or a frozen stare.
- Reflections: check windows, screens, glasses, and polished floors for inconsistent scenes.
- Background stability: watch for walls that shift, furniture that changes shape, or extras that vanish.
- Wardrobe and props: confirm colours and details match across shots.
- Light direction: confirm shadows fall consistently through the sequence.
- Text in frame: remove or replace any generated lettering.
- Audio sync: check lip movement against speech in every talking shot.
- Aspect ratios: verify vertical and square crops before delivery.
- First three seconds: if the hook does not land immediately, reorder rather than re-render.
A realistic timeline for a thirty-second explainer
A competent solo operator can produce a thirty-second product explainer across roughly six to ten hours of focused work, assuming the script already exists.
- Script to shot list: 45–60 minutes.
- Visual bible and reference stills: 1–2 hours.
- Generation and selection: 2–4 hours, the largest block.
- Voiceover and sound design: 1 hour.
- Edit, grade, and graphics: 1–2 hours.
- Quality checks and exports: 30 minutes.
The distribution matters more than the total. Budgets that collapse almost always under-fund reference stills and over-fund generation. A better bible reduces the number of takes you need by a wide margin, and it makes the second video in the same series dramatically cheaper than the first.
Frequently asked questions
Do I need to disclose that a video was made with AI?
Requirements vary by market and platform. Many advertising standards bodies expect any synthetic depiction of a real person to be disclosed, and several platforms require labelling of realistic synthetic media. The safe default is a short, unobtrusive disclosure in the description, plus on-screen text if a real person's likeness is simulated. Check the rules for each destination before you publish, because a single platform's policy is not a global standard.
Can generated product videos replace a real shoot entirely?
For concept, lifestyle, and abstract sequences, often yes. For a product that must be shown with engineering accuracy — precise dimensions, exact materials, a functioning interface — hybrid is usually better. Shoot or render the product itself, then generate the human and environmental footage around it. The audience will not notice the seam if the lighting and grade match.
How do I keep a character consistent across many shots?
Use a locked reference image, an identical wardrobe description copied verbatim between prompts, and the same lighting and lens language every time. Then cut into the character rather than holding on them for long stretches. Frequent cuts reduce the screen time over which a flaw can be noticed. If a shot absolutely must hold for six seconds on a face, plan to generate it later, after your references are proven on easier shots.
What clip length should I aim for?
Generate shorter than you think you need. Three to five seconds per shot is a comfortable working range for most engines and gives the editor room to trim. If a beat needs to last six seconds, generate two three-second versions and choose the stronger one, or split the beat into two shots with a cut in the middle.
How many generations should I plan for?
Assume a ten-to-one ratio between generations and usable clips for hero shots, and roughly three-to-one for simple B-roll. Anything better means your prompts and references are already well tuned. Track this ratio per project; when it climbs, the problem is almost always the references, not the model.
Is vertical video worth producing separately?
Yes, but do not crop a horizontal master and call it done. Generate or reframe with the vertical composition in mind, keep faces in the upper third, and regenerate any shot where the subject falls outside the safe area. Recomposition always beats a hard crop, particularly on shots with text or hands near the frame edge.
Should I invest in better prompts or better references first?
References, without question. A strong reference image with a plain prompt outperforms a beautifully written prompt with no reference in almost every test. Improve references until the base look is stable, and only then start refining the language of your prompts.
How do I keep a series visually coherent across multiple videos?
Freeze the bible. Keep the same palette, lighting rule, lens rule, and grade rule across episodes, and reuse the same voice and music family. Viewers recognise a series through texture and rhythm before they recognise the subject, so changing the grade between videos breaks the connection faster than changing the product.
The takeaway
Text-to-film works when you treat it as production rather than prompting. Write shots, not copy. Lock a visual bible before you render anything. Feed reference images, assemble prompts from reusable blocks, and mix engines by shot type instead of hunting for one perfect model. Then spend your remaining attention on sound and the edit, because that is where an assembly of clips becomes an advertisement.
The teams getting the best results from these tools are not the ones with the cleverest prompts. They are the ones with the tightest pre-production and the most disciplined quality checks. Build those two habits, keep a log of what worked, and the models will do the rest.


