Text-to-video has stopped being a novelty and started being a production line. What used to require a location scout, a camera crew, a cast, and a week of editing can now begin with a paragraph of copy and end with a finished vertical ad by the afternoon. But the gap between "a model generated something" and "this ad sells" is enormous, and it is almost never closed by the model itself. It is closed by the workflow around it.
This guide walks through that workflow end to end: how to write a script that survives generation, how to convert it into a shot list, how to prompt for specific visual outcomes, how to keep characters and products consistent across shots, how to edit and test, and where teams usually break their own results.
Why text-to-video reshaped ad production
Advertising has always been a race between iteration speed and media spend. A traditional shoot produces one or two usable cuts; testing ten emotional angles meant ten budgets. Generative video flips that equation. Once a concept exists as a script and a small set of reference assets, producing eight variants is a marginal cost rather than a second production.
Three technical shifts made this practical:
- Semantic understanding of long prompts. Modern language and diffusion systems parse compound instructions — subject, wardrobe, action, camera move, lighting, mood — instead of latching onto a single keyword.
- Temporal stability. Flicker, melting faces, and morphing hands used to make generated footage unusable beyond two seconds. Current models hold a subject together long enough for a real edit.
- Controllable cinematography. Camera paths, focal length hints, and motion strength can now be steered, which means generated shots can match the grammar of an existing brand campaign.
The result is a pipeline where the creative bottleneck moves from production to decision-making. You no longer run out of footage. You run out of clear ideas about what to test.
The workflow at a glance
A reliable AI ad pipeline has seven stages, and skipping any of them shows up on screen:
- Brief and promise. One sentence describing what the viewer should believe after three seconds.
- Script. A short narrative written for the ear, not the page.
- Shot list. The script decomposed into 6–14 discrete beats with durations.
- Asset bible. Reference images, character descriptions, product angles, color palette.
- Generation. Shot-by-shot prompting with a consistent seed and style token.
- Assembly. Editing, sound design, captions, aspect-ratio variants.
- Testing. Paid or organic distribution with structured variant comparison.
Most beginners collapse stages 1–4 into "write a prompt" and then wonder why the output looks generic. The prompt is the last mile, not the first.
Writing scripts that survive generation
The three-second contract
Every ad makes a promise in the first three seconds and pays it off before the end. On social feeds, the promise has to be visual, because sound is often off. That means your opening shot must communicate the tension without dialogue: a before/after, an impossible object, a scale mismatch, a human reaction.
Write the hook as a shot, not a sentence. "Someone realizes their routine is broken" is a script line. "Close-up of a hand reaching for a coffee cup that shatters into confetti" is a shot. Only the second one is generatable.
Keep sentences short and concrete
Generative models handle concrete nouns and visible actions well. They handle abstractions poorly. Compare:
- Weak: "She feels more confident about her finances."
- Strong: "She closes a laptop, stands up, and walks through a bright open-plan office, shoulders back."
A useful rule: if you cannot photograph it, you cannot generate it. Rewrite abstractions into behavior, wardrobe, and environment.
Structure the script as beats
A 30-second ad rarely needs more than four beats: hook, problem, turn, payoff. A 15-second cut needs two or three. Write each beat as a single line, then expand each line into one or two shots. When you are done, you should have a shot list that reads like a storyboard in text form.
Prompt engineering for ad shots
Prompting for video is closer to directing than to writing. A strong shot prompt has five components, and weakening any one of them shifts the output.
The five components
- Subject — who or what, with specific detail (age range, wardrobe texture, material).
- Action — a single continuous motion, not a sequence.
- Camera — shot size plus movement (slow dolly in, handheld follow, static macro).
- Light and environment — time of day, practical sources, weather, surface textures.
- Style and grade — film stock feel, color temperature, lens character, realism level.
A shot prompt might read: Macro shot of a matte black wireless earbud rotating slowly on a wet slate surface, water beading and rolling off, cool window light from the left, shallow depth of field, cinematic realistic grade, slow push in.
Notice that nothing in that prompt is a metaphor. Everything is a physical fact the renderer can resolve.
One action per shot
Models fail most often when asked to perform a sequence. "He picks up the phone, reads a message, laughs, and calls someone" is four shots, not one. Breaking it apart gives you control over pacing in the edit and dramatically improves fidelity.
Negative constraints matter
Most platforms accept an exclusion field. Use it for the artifacts that plague your specific subject: extra fingers, warped text, floating objects, lens flares, sudden zoom, distorted logos. Keep the list short — five to eight terms — because a long negative list dilutes attention.
Prompt libraries beat prompt genius
Once a shot works, save the prompt with the seed and reference image. Build a small internal library organized by shot type: product hero, human reaction, environment establishing, transition. New campaigns then start at 70% quality instead of zero.
Keeping characters, products, and scenes consistent
Consistency is the hardest problem in AI advertising, and it is where amateur results become obvious. A character whose jacket changes color between shots breaks the illusion instantly.
Build an asset bible
Before generating anything, assemble a folder containing:
- One or two reference images per recurring character
- Product photography from multiple angles
- A one-paragraph wardrobe and appearance description reused verbatim in every prompt
- A color palette with hex values for the grade
- A list of approved locations with 2–3 reference stills each
Reusing identical descriptive text across prompts is not laziness; it is the mechanism that keeps the model anchored.
Use seeds and image-to-video
The single biggest consistency upgrade is starting from a still. Generate or photograph a keyframe, then animate it. Character drift drops sharply because the model is interpolating from a fixed starting state rather than inventing one. Lock the seed where the platform allows it, and change only one prompt variable at a time when iterating.
Plan around the limitations
If a shot demands three people interacting with a product and precise hand contact, consider shooting that one plate practically and generating everything else. A hybrid pipeline is not a compromise; it is a professional choice. The audience never knows which shots were generated, only whether the whole thing feels coherent.
Directing camera, pacing, and sound
Camera language that reads on small screens
Vertical video rewards three moves: slow push in on a face or product, lateral reveal, and whip-pan transitions. Wide establishing shots are usually wasted on a phone. Shoot tighter than instinct suggests, and let motion carry the energy instead of cutting faster.
Cut on motion
Generated clips rarely have a clean beginning and end. Instead of fighting that, cut mid-motion. A hand entering frame in shot A and completing the gesture in shot B is invisible to the viewer and hides seams. Match direction of movement across cuts to keep continuity.
Sound carries generated footage
Sound design is the fastest way to make AI footage feel expensive. Layer three things: a bed (music or ambience), foreground foley tied to visible action, and a voiceover or caption rhythm. A soft whoosh on every cut, a subtle click when a product closes, and a low-pass filter on the music under the voiceover do more than any prompt change.
If dialogue is involved, decide early whether you are generating lip-sync or dubbing an on-camera-free voiceover. Voiceover with reaction shots avoids the uncanny valley entirely and is usually the safer route for direct-response ads.
Editing, captions, and format variants
The 4-3-2 cutting rule
From a 30-second master, cut a 15-second version and a 6-second bumper. Then export each in three aspect ratios and two caption styles. That is twelve deliverables from one shoot, and it is the reason AI production pays off.
Captions are not optional
Most viewers start muted. Burn in captions with a consistent font, high contrast, and a keyword emphasis style. Keep them inside safe areas so platform UI does not cover the payoff line.
Grade for cohesion
Apply one look across every shot — a subtle film emulation, a slight lift in the shadows, a consistent skin-tone bias. A unified grade can rescue footage generated across different sessions or even different models.
Testing and iteration without guesswork
Isolate variables
A useful variant test changes one thing: the hook, the payoff, the music, or the caption style. Changing three variables at once produces a winner you cannot explain and therefore cannot repeat.
Read the right metrics
- Hook rate (3-second view-through) measures the opening shot.
- Hold rate measures pacing and sound.
- Click-through rate measures the payoff and call to action.
- Conversion rate measures whether the ad kept the product's promise.
If the hook rate is strong and the hold rate collapses, the problem is in the middle, not the intro. Diagnosing by stage prevents random reshoots.
Recycle winning shots, not winning ads
Keep a shot-level library of what performed. A product rotation that held attention in one campaign is an asset you can drop into the next one. Over time, your library becomes the real competitive advantage — not access to any particular model.
Common mistakes and how to avoid them
- Prompting a sequence instead of a shot. Split multi-action descriptions into separate generations.
- Ignoring the storyboard. Generating before the shot list exists guarantees inconsistency.
- Chasing photoreal faces when hands and products would be safer. Choose your battles based on what the model does reliably.
- Over-stylizing. Heavy aesthetic filters date quickly and make product color grading impossible to match.
- Skipping sound. Silent AI footage reads as a demo reel, not an ad.
- No QA pass. Check hands, text, logos, reflections, and continuity on every single shot before assembly.
Choosing tools and building the pipeline
When evaluating platforms, weight these factors over raw demo quality:
- Shot-length reliability at the resolution you actually ship
- Image-to-video support, which drives consistency
- Motion and camera controls that are steerable, not random
- Commercial usage terms that fit your distribution and client work
- Export options including alpha channels, frame rates, and vertical crops
- Batch or API access if you produce more than a handful of ads per week
A pragmatic stack is usually hybrid: one strong model for human and environmental shots, one for product macro work, a standard editor for assembly, and a sound library. Avoid rebuilding your pipeline around every new release; rebuild only when a tool solves a bottleneck you have already measured.
Frequently asked questions
How long should each generated clip be?
Three to six seconds is the practical sweet spot. Longer clips tend to drift, and you rarely need more than a few seconds per beat in a short ad.
Can AI ads work for regulated industries?
Often yes for awareness and explainer content, with legal review of claims. Regulated products usually require accurate depiction of packaging and labeling, so plan on practical inserts for anything compliance-sensitive.
Do I need editing skills?
You need taste more than technical skill. Knowing when a shot is unusable and when a cut works is the deciding factor, and that transfers directly from traditional editing.
How many shots does a 30-second ad need?
Roughly ten to fourteen, including transitions. Fewer, longer shots feel cinematic; more, shorter shots feel energetic and perform better on social feeds.
What is the fastest way to improve output quality?
Move to image-to-video with locked reference stills and a consistent style token. This single change resolves most consistency complaints.
Should the script or the visuals come first?
Script first, always. The shot list derived from it is what makes generation predictable, and predictable generation is what makes iteration affordable.
Getting from text to screen reliably
The secret of effective AI advertising is not a prompt trick or a hidden model. It is discipline: a script written in photographable language, a shot list that treats each beat as one action, an asset bible that anchors characters and products, a grade and sound design that unify everything, and a testing loop that changes one variable at a time. Teams that follow that sequence produce ads that look intentional. Teams that skip it produce footage that looks generated — which is the one outcome an ad can never afford.


