Start Free Now
Limited Time Offer: Get 50% OFF Starter & Basic Yearly Plans 🎉

AI Video Synthesis Workflow: Luma and Veo in Practice

Sep 23, 2026

Why AI video synthesis finally fits real production pipelines

A few years ago, generating a moving image from a sentence was a novelty. Today it is a routine step in commercial work. Advertising agencies generate concept animatics before a shoot. Independent filmmakers build proof-of-concept sequences they could never afford to stage. Product teams spin up loopable hero clips for landing pages in an afternoon. Documentaries use synthesized inserts to cover gaps in archival material.

The reason for that shift is not simply that the images look better. It is that the tools became controllable. Early models produced a single impressive shot and then collapsed the moment you needed a second one that matched. Modern systems let you specify camera movement, hold a character's face steady across six cuts, and re-render a take without losing the rest of the sequence.

This guide is written for people who already understand the basics and want a working process. It compares two broad families of video models — cinematic-control systems in the spirit of Luma's Dream Machine line, and physically-grounded systems in the spirit of Google's Veo line — and then walks through a complete production workflow: prompting, shot design, continuity, finishing, quality control, and planning.

If you take one idea from this article, take this: model choice matters far less than pipeline discipline. A well-structured workflow on a mid-tier model beats random prompting on the best one available.

What actually happens inside a video model

Understanding the mechanics at a conceptual level changes how you prompt. You do not need to read research papers, but you do need to know what the model is optimizing for, because that determines where it fails.

Temporal coherence versus spatial fidelity

Every video model balances two competing goals. Spatial fidelity is how convincing a single frame looks in isolation. Temporal coherence is how well consecutive frames agree with each other.

A model that leans toward spatial fidelity produces gorgeous stills that shimmer, warp, and boil when played back. A model that leans toward temporal coherence produces smooth motion with softer detail and less textural richness. The two families described in this article sit at different points on that spectrum, and the trade-off shows up directly in your results: high-detail fashion beauty shots versus long unbroken camera moves through a physical space.

Conditioning: text, image, video, and structural guidance

Nearly all modern systems accept more than a text prompt. The common conditioning channels are:

  • Text — the base instruction, useful for subject, action, mood, and style.
  • First frame — an image that anchors the opening composition. This is the single most powerful control you have.
  • Last frame — pins the ending, which makes transitions between shots far easier to cut.
  • Reference images — portraits, costumes, or locations that the model reuses across generations for identity consistency.
  • Motion or depth guidance — extracted from a real clip or a 3D scene, applied to a generated one.

A prompt that only uses text is leaving most of the available control unused. Professional results almost always come from stacking several channels at once.

Why physics and optics became the new battleground

Once subject identity and frame stability were largely solved, differentiation moved to physical plausibility: how liquids pour, how cloth folds, how light refracts through glass, how a hand interacts with a solid surface. The newer generation of models has absorbed a great deal of implicit physical knowledge, which is why they can render believable collisions and reflections without explicit instruction. Where they still struggle is with compound interactions — a person handling a small object while walking through a crowded space, for example. Design shots that avoid stacking three difficult interactions in one take.

Choosing a model family: cinematic control or physical realism

The practical question is not "which model is best" but "which model is best for this shot, in this sequence, at this stage of the edit."

Where cinematic-control models excel

Models built around director-style control tend to offer granular adjustment over camera behavior and visual treatment. You can specify a slow dolly-in, an orbit, a crane rise, or a handheld drift, and the render respects it. They generally respond well to cinematic reference language: anamorphic flare, shallow depth of field, 35mm grain. They are strong choices for:

  • Title sequences and brand films where mood dominates action
  • Product hero shots that need precise camera choreography
  • Fashion and beauty work where lighting and skin detail matter most
  • Music-video style montages built from many short, stylized clips

Where physically-grounded models excel

Systems optimized for physical plausibility produce motion that reads as real. Wide shots hold together. Camera moves through complex environments feel weighty rather than gliding unnaturally. Characters walk with believable gait and interact with sets convincingly. They are strong choices for:

  • Narrative scenes with dialogue-adjacent blocking
  • Establishing shots and environment reveals
  • Action beats where weight and impact sell the moment
  • Anything that will be intercut with live-action footage

A simple decision table

Shot need Better fit
Precise camera move Cinematic-control family
Character walks through a crowd Physically-grounded family
Stylized product beauty Cinematic-control family
Insert shot matched to live action Physically-grounded family
Rapid montage of short clips Either — optimize for speed
Long unbroken take Physically-grounded family

The workflow that produces the best results uses both. Generate the stylized insert on one, the wide environmental shot on the other, and cut them together. Audiences do not know or care which model made which shot — they only register whether the sequence holds.

Prompting for shot-level control

A video prompt is not a description of a scene. It is a set of instructions for a camera crew and an art department that will only read it once.

Build a shot grammar

Write prompts in a consistent order so you can compare takes fairly. A structure that works reliably:

  1. Shot size and framing — wide, medium, close-up, over-the-shoulder
  2. Subject — who or what, described with two or three specific traits
  3. Action — one clear verb phrase, present tense
  4. Setting — where, with a time-of-day cue
  5. Camera — movement and lens character
  6. Light — source, quality, direction
  7. Look — grade, texture, film reference

Compare this to a typical first attempt, which is usually a poetic paragraph with no camera information at all. The structured version is shorter and far more controllable.

Camera and lens language that models understand

Most systems have absorbed standard cinematography vocabulary. Useful phrasing includes: slow dolly in, gentle push past the subject, static locked-off frame, slow orbit left, crane rise, handheld follow, rack focus from foreground to background, wide-angle distortion, telephoto compression, shallow depth of field with creamy bokeh.

One camera instruction per generation. Asking for a dolly-in that becomes an orbit which then cranes up will produce mush or, worse, a jump cut inside a single clip.

Lighting and color as control signals

Lighting descriptions do double duty: they set mood and they reduce ambiguity about geometry. "Soft window light from camera left, cool shadows, warm highlight on the cheek" tells the model where the light source is and how surfaces should behave. Vague terms like "beautiful lighting" tell it nothing.

What to leave out

Omit anything you cannot verify in a still frame. Internal emotional states, backstory, and abstract concepts do not survive translation into pixels. "She feels conflicted about leaving" becomes visible only through specific staging: a pause at the doorway, a hand on the frame, a glance back. Describe the staging, not the feeling.

Also resist the temptation to over-prompt. Beyond a certain length, additional adjectives start competing with each other. If you need six sentences, consider whether this should be two shots instead.

Building a repeatable multi-shot workflow

The difference between hobby output and professional output is that professionals generate shots in a predictable order with known dependencies.

Pre-production: decide the look before you generate

Before touching a video model, lock three things: aspect ratio, target runtime, and a visual reference board. Collect ten to fifteen stills that represent the intended grade, lens character, and production design. This board will keep you consistent across days of work and across collaborators.

Then break the sequence into shots on paper. For a thirty-second piece, ten to fourteen shots is typical. Assign each shot a purpose: establish, introduce, detail, transition, resolve. Shots without a purpose get cut in the edit anyway.

Keyframes first, motion second

Generating still keyframes before animating them is the single highest-leverage habit in this workflow. A still image costs a fraction of a video generation, renders in seconds, and can be revised indefinitely. Iterate on composition there, where feedback is fast, rather than on expensive moving footage.

Workflow:

  1. Generate or shoot a first-frame image for every shot in the sequence.
  2. Place all of them on a timeline as a stills animatic with rough timing.
  3. Watch it. Fix the story before you spend time on motion.
  4. Only then send each approved frame into the video model.

This approach routinely saves more than half the total generation effort, because you eliminate shots you would otherwise have animated and discarded.

Image-to-video versus text-to-video

Text-to-video is best for exploration and for shots where the model's own composition instincts are an asset — landscapes, abstract textures, atmospheric transitions. Image-to-video is best for anything that must match an approved frame: characters, product angles, established locations.

A hybrid pattern works well: use text-to-video to discover a look, freeze your favorite result as a still, then regenerate the sequence through image-to-video for control.

Consistency across shots: character, props, and location

Audiences forgive imperfect physics. They never forgive a character whose face changes between cuts.

Character consistency

Use reference images rather than description. Most modern systems accept one or more portrait references and will carry facial structure, hair, and wardrobe across generations. Keep a dedicated reference folder per character with a neutral expression, a three-quarter view, and one full-body shot.

If a model still drifts, narrow the variables: lock camera height, lock focal length, and avoid extreme angles until the identity is stable. Once you have a reliable take, reuse its first frame as the reference for the next shot in the same scene.

Prop and wardrobe continuity

Small details break continuity most often: a jacket that changes color, a cup that switches hands, a logo that flips. Handle these by keeping a written continuity sheet — one line per shot listing wardrobe, key props, and their positions. It sounds bureaucratic and it saves entire reshoots.

Location and lighting continuity

Generate one hero wide shot of each location first and treat it as canon. All closer shots in that space should inherit its light direction, color temperature, and time of day. Changing the light direction between a wide and a close-up is the fastest way to make a sequence feel assembled from unrelated parts.

Audio, editing, and finishing

Generated footage is raw material, not a finished piece. The finishing pass is where most of the perceived quality is added.

Sound design

Silent AI footage feels artificial no matter how good the image is. Add ambience, foley, and music before you judge a cut. A room tone bed and a few well-placed footsteps will make a synthetic shot read as real more effectively than another hour of regeneration.

For voice, generated narration or dialogue should be recorded or synthesized at a consistent level and then compressed slightly before mixing. Dialogue that is much louder than the music draws attention to the fact that it was added afterward.

The edit

Cut on motion. If a camera is pushing in, cut during the push; if a subject is turning, cut mid-turn. Motion-matching hides imperfections and gives sequences energy that static cuts cannot.

Keep generated clips short — two to four seconds is usually plenty. Long generated takes expose every inconsistency, while a sequence of short ones feels intentional and controlled.

Cleanup and upscaling

A finishing pass often includes:

  • Frame interpolation to a consistent frame rate across mixed sources
  • Temporal denoise to remove flicker in shadows and flat areas
  • Face restoration, applied gently — heavy settings produce a plastic look
  • Upscaling to delivery resolution
  • A unifying grade applied across all shots to bind them together

Applying one consistent grade at the end is essential when shots come from different models. It is the cheapest way to make a mixed-origin sequence feel like one film.

Quality control checklist and common mistakes

Run every sequence through the same checklist before delivery. Consistency of process catches errors that consistency of model never will.

Before you generate:

  • Aspect ratio and frame rate locked and documented
  • Reference board complete
  • Shot list with purposes assigned
  • Stills animatic approved

Before you deliver:

  • Watch the full sequence at normal speed with sound on
  • Watch it once at half speed looking only at hands and faces
  • Watch it muted to check whether the story reads visually
  • Check continuity sheet against final frames
  • Verify no unintended text, watermarks, or garbled signage
  • Confirm color and loudness match the delivery specification

Common mistakes:

  1. Prompting a paragraph instead of a shot. Poetic prompts produce unpredictable results. Structure wins.
  2. Anchoring on text only. A first-frame image removes most ambiguity for free.
  3. Changing two variables at a time. Change one thing per iteration or you will not know what fixed the problem.
  4. Assuming a rejection means the idea is bad. Often the same idea works on a different model or from a different starting frame.
  5. Skimping on sound. Audio quality is the fastest route to perceived professionalism.
  6. Forgetting transitions. Shots that work individually can still refuse to cut together. Generate last frames with the next shot's composition in mind.

Planning time, cost, and hardware

Treat generation the way you would treat a shoot day: budget it, schedule it, and keep a shot-by-shot log.

Three variables trade off against each other — quality, iteration count, and turnaround. You can have two. A quick-turnaround deliverable uses fewer iterations and accepts a slightly less polished look. A prestige piece buys iterations with time.

Practical planning habits:

  • Estimate three to five generations per final second of footage for controlled work
  • Batch similar shots together so your prompts stay in the same mental space
  • Keep every prompt and every accepted frame in a versioned folder
  • Log which model and which settings produced each accepted shot so you can reproduce it later
  • Reserve a fixed portion of your time budget for the finishing pass, which is easy to forget and always takes longer than planned

For hardware, cloud generation removes local constraints but introduces queue latency during peak hours. Local generation demands a capable GPU but gives unlimited iteration without per-run accounting. Hybrid workflows — local for exploration, cloud for final high-resolution renders — are common in professional contexts.

FAQ

Do I need two different models to get good results?
No, but a single model has a defined comfort zone. Understanding where yours is strong — and designing shots that play to it — matters more than adding tools.

How long should a generated clip be?
Two to four seconds. Short clips hide imperfections, cut together cleanly, and are far easier to regenerate after a retake request.

Why does my character's face change between shots?
Almost always because you relied on text description rather than reference images, or because you changed camera distance dramatically. Lock references and keep framing similar within a scene.

Is image-to-video always better than text-to-video?
No. Text-to-video is better for discovery and atmosphere. Use image-to-video once you have an approved frame you need to match.

How many attempts should one shot take?
Three to five is normal for controlled work. Ten or more usually means the prompt is overloaded or the shot concept is too complex for one take.

What is the most common reason a sequence feels artificial?
Missing sound design and inconsistent color. Both are finishing problems, not generation problems, and both are cheap to fix.

Can generated footage be intercut with live action?
Yes, and it works surprisingly well when you match grain, black level, and lens character in the finishing pass. Inconsistent texture between sources is what gives the trick away.

Where should a beginner start?
With a stills animatic. Build the sequence as images first, learn what compositions hold the story, and only then start animating. That single habit prevents most wasted effort.

The technology will keep changing. Pipelines and habits persist. Build a process you can run on any model, and every new release becomes an upgrade rather than a rewrite.

Alexander

Alexander