Text-to-video generation has quietly crossed the line between novelty and tool. A steady run of incremental releases has produced models that hold a subject together across a camera move, respect a reference image, and deliver a clip you can drop into a timeline without rebuilding it frame by frame. Pika 2.3 and Veo 3.5 are two useful examples of that shift, and they represent two different philosophies: Pika leans toward fine-grained motion control and rapid iteration, while Veo leans toward narrative length, scale, and turnaround speed. Neither wins everywhere. The productive question is not which model is better in the abstract, but which shot you are trying to land — and what workflow gets you there without ten rounds of guesswork.
Why text-to-video crossed a practical threshold
The old complaints about generated video were consistent and fair. Subjects melted. Hands fused into mittens. A character walking toward the camera would drift sideways for no reason. Camera language in a prompt was treated as decoration rather than instruction. The result was footage that looked impressive for three seconds and unusable for anything longer.
What changed is not one dramatic breakthrough but a cluster of improvements arriving together. Motion coherence improved, meaning an object keeps its shape and trajectory across a pan or a push-in. Camera interpretation improved, meaning phrases like slow dolly in or handheld tracking shot now produce measurably different footage. Reference conditioning improved, meaning an uploaded frame or character sheet actually anchors wardrobe and lighting instead of being vaguely echoed. Clip length increased enough to cover a real beat rather than a fragment.
Put those together and the bottleneck moves. Generation is no longer the slow part of the process; editorial judgment is. You can produce six variations of a shot before lunch and spend the afternoon deciding which one serves the story. That is a completely different job than coaxing a single usable frame out of a stubborn model, and it changes how you should plan your time.
It also changes the economics of previsualization. Directors and marketers can now build an animatic that looks close to the final product for a fraction of what a live shoot would cost, then use it to test pacing, framing, and tone before committing real resources. The generated footage is not a replacement for production — it is a rehearsal space.
How to evaluate a model before you commit to it
Model announcements tend to emphasize peak quality on cherry-picked prompts. That is the least useful signal for day-to-day work. When you are choosing a model for a project, evaluate it on the dimensions that actually affect your schedule.
Motion coherence. Generate the same prompt that requires sustained movement — someone crossing a room, a car turning a corner, fabric moving in wind — and watch the fifth second, not the first. Most models look strong at the start and degrade as the clip continues.
Prompt adherence. Write a prompt with four explicit constraints: subject, action, camera move, and lighting. Count how many survive into the output. Models that honor three out of four consistently are far more workable than models that produce beautiful footage that ignores your instructions.
Camera vocabulary range. Test whether distinct camera terms produce distinct results. If every camera phrase yields the same slow drift, you cannot direct the shot, only suggest it.
Reference conditioning. Upload a character image and a location image together. Check whether both influence the output or whether one dominates. Consistency across a multi-shot sequence depends entirely on this behavior.
Duration and resolution per generation. Short clips are fine if you can chain them, but chaining introduces continuity risk. Know your practical maximum before you build a shot list around it.
Speed and pricing structure. Time-to-first-frame matters when you iterate heavily. So does how the platform meters usage — per second of output, per generation, or through a subscription allowance. A model that is slightly weaker but three times faster often produces better final results because you can afford more attempts.
Run this test suite on your own material, not on someone else's demo. Twenty minutes of structured testing tells you more than a week of reading comparisons.
Prompt architecture: writing for motion, not for stills
The most common mistake in text-to-video is writing image prompts and expecting video. Image prompts describe a frozen composition. Video prompts have to describe a trajectory, because the model needs to know where things start, where they go, and how the camera relates to both.
A reliable prompt skeleton has six slots, filled in this order:
- Subject — who or what, with two or three defining visual details.
- Action — a single dominant verb describing change over time.
- Camera — shot size plus one movement.
- Environment — location, weather, time of day, background activity.
- Light — direction, quality, and color temperature.
- Style — film stock, lens character, animation style, or reference era.
Two rules keep this skeleton from collapsing. First, one dominant action per shot. If your prompt contains someone walking, turning, opening a door, and picking up a phone, the model will average these into a vague shuffle. Split it into multiple shots instead. Second, order matters more than you expect — put the elements you care most about early, and repeat the single most important one at the end as reinforcement.
Verb choice is also load-bearing. Slow, deliberate verbs produce slow, deliberate motion. Energetic verbs produce energetic motion. Adjectives like cinematic or beautiful add almost nothing; they are style noise that the model cannot convert into a decision. Replace them with concrete descriptors: shallow depth of field, tungsten practical lights, 35mm grain, overcast daylight.
For dialogue or on-screen text, set expectations low. Rendering readable text inside generated footage remains unreliable, and lip-sync quality varies wildly with head angle and speaking pace. If a shot needs precise words, generate the visual clean and add text or voice in post.
Camera control vocabulary that changes the output
Once models began interpreting camera language reliably, prompting became closer to directing. The terms below produce distinguishable results across most modern models, including Pika 2.3 and Veo 3.5, though the exact rendering differs.
- Static / locked-off — no movement. Underrated for dialogue and product detail shots.
- Push in / dolly in — camera moves toward the subject. Good for realization and rising tension.
- Pull out / dolly out — camera retreats. Good for reveals and endings.
- Truck left / right — lateral movement parallel to the subject.
- Crane up / down — vertical movement, useful for scale.
- Orbit — circling the subject. Powerful but stresses consistency.
- Handheld tracking — follows a moving subject with slight instability.
- Whip pan — fast rotation, best used as a transition.
- Rack focus — shifts focus between foreground and background planes.
Two practical caveats. First, combine shot size with movement rather than stacking multiple movements: medium shot, slow push in works; medium shot, push in then orbit then crane does not. Second, movement costs coherence. The more the camera travels, the more chances the subject will warp, especially at the edges of the frame. If a shot keeps failing, reduce camera movement before you rewrite the subject description.
Lens and format language also carries weight. Specifying a wide lens, anamorphic flare, or a 2.39:1 framing shapes the output in ways that feel more natural than describing mood. Aspect ratio deserves particular attention, because generating in the wrong ratio and cropping later silently destroys composition.
Consistency across shots: characters, wardrobe, and locations
Single clips are easy. Sequences are where projects live or die. A character who changes jacket color between shots breaks the illusion faster than any rendering artifact.
Build reference assets before you generate anything. Three items pay for themselves immediately:
A character sheet. One clean, front-facing, neutral-light image of each recurring character, plus a second angle if the model accepts multiple references. Include wardrobe details explicitly: color, material, distinguishing features.
Location plates. One wide establishing frame per location, used as a reference for every shot set there. This locks architecture, palette, and light direction.
A style anchor. A single frame that defines the grade, grain, and contrast you want across the whole sequence. Apply it as a reference even on shots where it seems unnecessary.
With those assets in place, keep your descriptive language repeatable. Write the character description once and paste the identical phrasing into every prompt. Small rewordings — changing a blue wool coat into a navy jacket — will produce visible drift. Boring consistency in prompts creates consistency on screen.
Reality also has a role. If a sequence requires heavy continuity, storyboard it as a set of locked-off or lightly-moving shots rather than a chain of dramatic camera moves. Save the ambitious camera work for the moments where the story earns it.
A repeatable end-to-end workflow
The workflow below assumes a short piece: thirty to ninety seconds, five to fifteen shots. Scale it proportionally for longer work.
Step 1: Script to shot list
Write the piece as prose first, then break it into beats. Each beat becomes a shot with a stated purpose: establish, advance, react, transition, resolve. If you cannot state a shot's purpose in a few words, cut it. Convert each shot into the six-slot prompt skeleton and note which reference assets it needs.
Step 2: Keyframes and references
Generate or source a still for every shot before generating video. Stills are cheap and fast. They let you fix composition, wardrobe, and lighting while changes cost nothing. When a still is right, use it as an image-to-video start frame. This single habit eliminates more wasted generation than any prompting trick.
Step 3: Generation passes
Work in two passes. The first pass is low-resolution exploration: two or three variations per shot with loose prompts, purely to find which interpretation reads best. The second pass is the real one: lock the prompt, attach references, and generate at final quality, usually with two or three variations for safety.
Keep a simple log — prompt text, reference used, model, settings, result rating — because after twenty generations you will not remember which phrasing worked. A spreadsheet is enough.
Step 4: Assembly and finishing
Bring the clips into your editor and cut for rhythm before you fix anything else. Pacing solves many problems that look like quality problems. Then handle the small stuff: stabilize shaky shots, crop to a consistent ratio, unify color with a single grade or LUT, add grain or sharpening to hide resolution differences between models. Mixed-source sequences reveal themselves through inconsistent grain and contrast more than through anything else.
Step 5: Sound
Sound is not a final garnish. Diegetic sound — footsteps, room tone, traffic — sells generated footage enormously, because the eye forgives visual softness when the audio world is coherent. Generate or source ambience per location, add a music bed that matches the cut rhythm, and keep dialogue processed lightly. If a character speaks, consider showing them from behind or in profile to avoid lip-sync scrutiny.
Speed, cost, and quality trade-offs
Every project sits somewhere on a triangle of speed, cost, and quality, and pretending otherwise leads to blown schedules. Practical rules that hold up:
- Draft cheap, finish expensive. Iterate at the lowest quality that still tells you whether the shot works.
- Budget retries, not just shots. Assume two to three attempts per final shot. If your plan has no room for retries, it has no room for quality.
- Parallelize where the interface allows. Queue several generations and review them together rather than waiting on one at a time.
- Match resolution to delivery. Social vertical video rarely benefits from maximum resolution; a big-screen piece does.
- Prefer fewer, better shots. Ten confident shots beat twenty-five uncertain ones and cost less to finish.
Speed also affects creative ambition in a subtle way. When turnaround is fast, you experiment more and discover better ideas. When it is slow, you play safe. Factor that into your choice of tool, not just the per-second economics.
Common mistakes and their fixes
Overstuffed prompts. Four actions in one sentence produce mush. Fix: one dominant action per shot, split the rest.
Conflicting camera instructions. Push in and pull out in the same prompt cancel each other. Fix: one movement, chosen deliberately.
Ignoring aspect ratio. Generating widescreen and cropping to vertical destroys composition. Fix: set the delivery ratio before you write the prompt.
No style anchor. Each shot ends up looking like a different film. Fix: attach a reference frame and repeat identical style language.
Expecting reliable on-screen text. It rarely renders cleanly. Fix: generate a clean plate and add typography in the edit.
Chasing a failed shot with more words. If three revisions fail, the problem is usually camera movement or too many subjects. Fix: simplify geometry before rewriting description.
Treating audio as an afterthought. Silent generated footage feels artificial. Fix: plan ambience and music from the storyboard stage.
Decision guide: matching tools to shots
Different models suit different jobs, and a hybrid approach usually beats loyalty to one.
Pika 2.3 tends to reward deliberate motion and camera control. It is a strong choice for stylized sequences, short punchy social clips, and shots where you want to iterate quickly through framing ideas. Its reference handling makes it comfortable for consistent character work when paired with a character sheet.
Veo 3.5 tends to reward narrative continuity and scale. It suits longer beats, environments with many moving elements, and projects where turnaround speed matters as much as polish. It is often the better pick for establishing shots and sequences that need to hold together over time.
Open-source options such as Tencent Hunyuan Video are worth testing when you need local processing, unusual fine-tuning, or cost predictability at volume. They demand more setup and hardware attention, but they remove dependencies that matter for some teams.
A practical hybrid: block out the whole sequence in whichever model iterates fastest for you, then re-render hero shots in whichever model produces the strongest image. Keep the reference assets shared between them so continuity survives the switch.
FAQ
Do I need to know cinematography to get good results? It helps more than any prompting trick. Shot size, movement, and light direction are the vocabulary the models respond to. You can learn the basics in an afternoon and it will improve your output permanently.
How long should a single generated shot be? Short enough that nothing degrades. Many creators keep shots between three and eight seconds and build longer sequences through editing rather than generation. Test your model's tolerance and cut before it fails.
Why does my character change between shots? Almost always because the descriptive language changed or no reference image was attached. Fix the phrasing, attach a character sheet, and keep wardrobe wording identical across every prompt in the sequence.
Should I generate stills first or go straight to video? Stills first. It is faster, cheaper, and it forces you to solve composition before motion complicates the problem. Most wasted generation comes from skipping this step.
How many attempts should I plan per shot? Two to three for final quality, more for shots with heavy camera movement. If a shot needs ten attempts, simplify the prompt rather than expanding the retry budget.
Can generated video replace traditional production? For previsualization, social content, and stylized sequences, increasingly yes. For anything requiring precise performance, complex interaction, or legal clarity about likeness, treat it as a complement rather than a substitute.
What is the fastest way to improve? Keep a log of prompts and results, review it weekly, and copy what worked. Pattern recognition beats theory, and the patterns are personal to your subject matter.



