Start With the Shot, Not the Sentence
Most disappointing AI video results come from a simple mismatch: the person writing the prompt is thinking in sentences, while the model is thinking in shots. Luma Dream Machine and Luma 2.1 respond best when you describe a single, filmable moment with a clear subject, a defined action, a camera position, and a lighting condition. Everything else is decoration.
This is why two prompts of similar length can produce wildly different output. "A woman walks through a rainy city at night, cinematic, 4k" gives the model a mood but almost no staging. "Medium shot, a woman in a soaked trench coat walks toward camera along a wet Tokyo alley, neon signs reflecting in puddles, slow handheld push-in, shallow depth of field" gives it a plan. The second prompt is not longer for the sake of length — every clause removes a decision the model would otherwise make randomly.
This guide is a practical workflow for prompting Luma's video models. It covers prompt anatomy, camera language, realism cues, image-to-video techniques, continuity across shots, a repeatable testing loop, and the mistakes that quietly ruin otherwise good generations.
What Luma Actually Reads in a Prompt
Luma's models are trained on paired video and text, so they have learned associations between phrasing patterns and visual outcomes. That means word choice has consequences: "dolly in" is not a synonym for "zoom in," and "golden hour" produces different color science than "warm sunlight."
The seven slots of a reliable prompt
Think of every prompt as a set of slots. Skim a draft and check whether each is filled — or deliberately left empty.
- Subject — who or what, with one or two defining details.
- Action — a single continuous motion, not a sequence of events.
- Environment — location, weather, time of day, background activity.
- Camera — shot size, angle, lens feel, movement.
- Lighting — source, direction, quality, color.
- Style — film stock, era, genre, color grade, rendering look.
- Constraints — what to avoid, plus pacing and duration expectations.
A useful rule: if a clause does not change the picture, cut it. "Award-winning, masterpiece, trending, ultra-detailed" rarely changes the picture. "Shot on 16mm, slight grain, muted teal shadows" does.
Order and weighting
Luma tends to weight earlier tokens more heavily in text-to-video. Put the subject and the primary action in the first sentence. Stylistic language works better as a trailing clause than as an opening flourish. If a detail matters enormously — a red scarf, a specific lens — repeat it once in a different phrasing rather than stacking five adjectives on it.
Negative phrasing is weaker than positive phrasing
Saying "no camera shake" is less effective than saying "locked-off tripod shot." Saying "not crowded" is weaker than "empty street at dawn." Models follow instructions better when you describe the thing you want instead of the thing you do not.
Text-to-Video Versus Image-to-Video
Luma Dream Machine supports both text-to-video and image-to-video, and choosing the wrong mode is one of the most common sources of wasted effort.
Use text-to-video when discovery matters
Text-to-video is best for exploration: mood boards, concept tests, style sampling, and quick ideas where the exact composition is negotiable. Describe the shot and let the model surprise you. Generate four variations with small changes — different camera height, different time of day — and keep the one that reads clearly even at thumbnail size.
Use image-to-video when the composition is already decided
If a client approved a storyboard frame, or you have a still you love, start from the image. The model preserves composition and identity, and your prompt shifts from describing what exists to describing what happens: how the camera moves, how hair and fabric react to wind, how light changes.
Image input: portrait still, subject facing three-quarter left
Prompt: slow orbit right around subject, hair lifting in a light breeze,
window light shifting as clouds pass, 35mm lens, subtle film grain,
shallow depth of field, natural color grade
Keyframe chaining
For a two-shot sequence, generate the first clip, export a clean frame from the end, then use that frame as the starting image for the next clip. Match the motion direction across both prompts — if clip one ends with a leftward pan, clip two should continue leftward, not reverse. This simple habit makes sequences feel intentional rather than stitched.
Camera Language That Actually Moves
Camera vocabulary is the highest-leverage part of a Luma prompt because it is the least ambiguous. The model has a clear mapping between these terms and motion patterns.
The core moves and what they feel like
| Prompt phrase | Visual result | Best used for |
|---|---|---|
| static shot / locked-off | No camera movement | Product shots, portraits, dialogue |
| slow push in | Gradual approach | Building tension, focus on detail |
| pull back / reveal | Camera retreats | Revealing scale or context |
| tracking shot | Camera follows subject | Walking, running, vehicles |
| orbit / arc around | Circular movement | Hero shots, character intros |
| crane up | Vertical rise | Landscapes, endings |
| handheld follow | Slight instability | Documentary, urgency |
| aerial drone shot | High altitude | Establishing geography |
Combine at most two moves
"Slow push in while tilting up" reads cleanly. "Push in, orbit, crane up, then snap zoom" produces mush, because the model has to average conflicting trajectories within a very short runtime. Pick a primary move and, if needed, one subtle secondary adjustment.
Specify lens and framing alongside the move
Shot size and lens feel change how the movement reads. A wide 24mm push in feels environmental; an 85mm push in feels intimate and compresses the background. Adding "35mm lens, shallow depth of field" tells the model how much of the frame should fall out of focus, which in turn controls how much attention the subject receives.
Pace words matter
"Slow," "steady," "deliberate," "snappy," and "frantic" all shift timing. When a clip feels too fast, the fix is often not a different prompt but a different pace adjective plus an explicit duration cue, such as "over five seconds, gradual."
Realism: Lighting, Texture, and Physics
Luma 2.1 is noticeably stronger than earlier generations at physically plausible motion — fabric that folds, water that splashes, dust that drifts. You can lean into that strength with specific material and lighting vocabulary.
Lighting descriptors that behave predictably
- Source: practical neon, fluorescent office, overcast sky, single tungsten lamp, car headlights.
- Direction: backlit, side-lit from the left, top-down, rim light on the shoulders.
- Quality: hard shadows, soft diffused wrap, dappled through leaves.
- Color: cool blue shadows with warm highlights, sodium-vapor orange, desaturated teal.
Combining source plus direction plus quality gives you more control than a single adjective like "cinematic." "Cinematic" is a grade, not a light.
Texture and material cues
Naming materials tells the model how surfaces should reflect and deform: brushed aluminum, matte ceramic, cracked leather, wet asphalt, brushed wool, frosted glass. Add micro-detail only where the camera will actually see it — close-ups benefit from skin texture and fabric weave; wide shots benefit from atmosphere and depth layers.
Where realism breaks and how to prompt around it
Hands, crowds, and on-screen text remain the fragile zones. Practical mitigations:
- Keep hands partially out of frame, in pockets, or holding an object with a simple silhouette.
- Replace crowds with "a few distant figures, silhouetted" rather than "busy crowd."
- Avoid legible signage in the prompt; if a sign is needed, keep it out of focus or describe it as abstract shapes.
- For any action involving tools or props, describe the prop, the grip, and the direction of motion in one clause.
Motion physics vocabulary
Words like "momentum," "heavy," "light," "settling," "trailing," and "swaying" influence how weight reads. "A heavy canvas curtain sways and settles" produces better motion than "a curtain moves," because it tells the model the mass of the object.
Multi-Image Fusion and Style Locking
When you can supply more than one reference, the goal is usually consistency: the same character, the same palette, the same world across several clips.
A reliable approach is to separate references by role. One image establishes identity, one establishes environment, one establishes the color and texture treatment. Then write a prompt that references those roles implicitly rather than listing them: describe the character, the location, and the grade in the same sentence structure you would use for text-to-video.
To keep a look consistent across a series, save a short "style block" — a fixed phrase of 15 to 25 words covering lens, grain, color, and lighting quality — and paste it at the end of every prompt in the series. Consistency across clips comes more from repeating identical style language than from repeating identical subject language.
Sequences, Loops, and Continuity
Single clips are easy to judge and hard to build with. Sequences are where prompting becomes a production skill.
Designing a three-shot sequence
Use a wide to establish, a medium to explain, and a close-up to emphasize. Write each prompt as a standalone shot, but keep three things identical across all three: the lighting direction, the color grade block, and the motion direction of the camera.
Loop-friendly prompts
For seamless loops, the camera must return to its starting position and the subject's motion must resolve. Prompts that say "continuous circular camera orbit returning to start position" or "rhythmic repeating motion" loop far better than prompts with narrative progression. Ambient subjects — steam, rain, traffic, flags, fire — loop almost effortlessly.
Planning for aspect ratio and duration
Vertical crops demand tighter framing; horizontal frames reward depth layers. Write your prompt for the aspect ratio you intend to deliver, because "wide establishing shot" in a 9:16 frame produces an awkward, empty image. Similarly, a prompt describing three separate beats will not fit into a five-second clip — cut it down to one beat.
A Repeatable Prompt Testing Workflow
The difference between hobbyist results and professional results is usually process, not vocabulary. A short loop beats a long prompt every time.
- Write the shot as a sentence. Who, doing what, where, and what the camera does. One line.
- Add the cinematic layer. Convert that line into slots: subject, action, environment, camera, lighting, style, constraints.
- Generate three variations. Change one variable each time — camera height, time of day, pace word. Never change three things at once or you learn nothing.
- Judge at delivery size. Watch at the size and aspect ratio the audience will see. Shots that look great full-screen often fail as feed thumbnails.
- Lock the best prompt. Copy it into a notes file with the date and a one-line description of what worked.
- Extend from the winner. Change the subject or location, keep the camera and style block, and generate the next shot in the sequence.
- Only then add complexity. Multi-image references, character locking, and long sequences should come after you have a working single-shot prompt.
Keep a personal library of prompts that produced clean results. Over time, that library becomes more valuable than any generic prompt list, because it encodes your own style and the specific quirks you have learned to work around.
Common Mistakes and How to Fix Them
| Symptom | Likely cause | Fix |
| --- | --- |
| Video feels generic and flat | Prompt is all mood, no staging | Add shot size, lens, and lighting direction |
| Camera motion is chaotic | Three or more competing moves | Reduce to one primary move |
| Cuts feel disconnected | Style language changed between clips | Reuse an identical style block |
| Face drifts between shots | No image reference for identity | Start from a still or reference image |
| Motion looks weightless | No material or mass cues | Add material words and motion physics terms |
| Prompt ignored in second half | Too many beats for the clip length | One beat per clip |
| Everything looks over-graded | Stacked style adjectives | Pick one grade, name it once |
Two more habits worth building: avoid stacking adjectives, and never leave a prompt untested because it "sounds good." The model does not judge prose, it interprets tokens.
Choosing the Right Tool for the Job
Luma Dream Machine is strong on atmospheric realism, natural camera motion, and image-to-video continuity. Other tools in the same family of generative video systems have different strengths — some excel at stylized motion, others at longer single takes or precise character control. A practical workflow often mixes them: prototype the mood in one model, refine the shot in another, and finish in an editor where you can trim, stabilize, interpolate frames, and color match.
The mistake is treating any model as the final step. Generated clips become usable footage only after editing — pacing, sound design, music, and grade do as much for perceived quality as the prompt did. If you are building a deliverable, budget time for post-production equal to the time you spent generating.
Also consider resolution and duration trade-offs early. Short clips are easier to make convincing, which is why so many strong AI video edits rely on cuts every two to four seconds. If a clip must run long, prompt for a single slow movement and let the edit do the rest.
FAQ
How long should a prompt be?
Most reliable prompts land between 30 and 70 words. Under 20 words leaves too many decisions to the model; over 100 words dilutes attention and increases the chance that later clauses are ignored.
Should I include style words like "cinematic" or "8k"?
Sparingly. Naming a concrete reference point — lens, film stock, lighting quality, color palette — is far more effective than abstract quality claims.
Why does my camera move ignore my instructions?
Usually because the prompt contains multiple conflicting moves, or because the move is buried after several descriptive clauses. Put the camera instruction early and keep it to one primary movement.
How do I keep a character consistent across clips?
Use an image reference for identity, repeat the same description wording for the character, and keep the style block identical. Consistency comes from repetition, not from adding detail.
What makes a clip loop cleanly?
Matching start and end camera positions, rhythmic subject motion, and ambient elements like steam, rain, or fire. Avoid narrative progression in loop prompts.
Do negative prompts help?
They help less than positive alternatives. Describe the state you want — "empty street," "locked-off tripod shot" — rather than the state you want to avoid.
How many generations should a single shot take?
Expect three to five for a good result and ten or more when you are learning a new subject type. Track what changed between attempts so each generation teaches you something.
Where to Go From Here
The shortest path to better AI video is not a longer prompt list — it is a tighter loop. Pick one shot you care about, fill the seven slots, generate three variations, and keep the one that reads clearly at delivery size. Then repeat that loop with one variable changed at a time.
Once single shots feel predictable, move to sequences: lock a style block, plan wide-medium-close coverage, and chain keyframes to preserve continuity. Add multi-image references only when a shot genuinely needs identity or environment locking.
Luma Dream Machine and Luma 2.1 reward specificity, physical plausibility, and restraint. Write like a camera operator describing the next setup, not like a marketer describing a finished film — and your generations will start looking like footage instead of guesses.

