Why AI Video Has Become an Ordinary Production Tool
Something shifted in the last few production cycles. AI video generation stopped being a novelty that agencies demonstrated on internal innovation days and became a line item in real schedules. Not because the tools are magical — they are not — but because they compress specific, expensive parts of the pipeline: look development, previsualisation, pickups that no longer require a full reshoot, and the endless small inserts that otherwise eat a day of crew time for eight seconds of screen time.
In the Netherlands, where crews are highly skilled and day rates reflect that, the compression is easy to justify. A commercial house in Amsterdam can pitch three distinct visual directions in the time it previously took to pitch one, because each direction can be previsualised with generated footage instead of described in a deck. A documentary team can visualise an archive scene that no longer exists, or a future scenario that cannot be filmed, without booking a camera package and a drone permit.
The interesting part is who adopts fastest. It is rarely the youngest editor chasing trends. It is experienced directors and producers who already know exactly what a shot needs to achieve and are willing to spend an hour testing whether a model can deliver it. Experience is the multiplier. The same tool produces forgettable results for someone who has not decided what the scene is about, and usable results for someone who has. Directing is still directing; the medium of capture changed.
This guide is deliberately neutral and workflow-first. It covers what current models genuinely do well, how to choose between them without chasing hype, how to build a pipeline that survives client review and broadcast delivery, and where the whole approach still falls over.
What These Tools Are Genuinely Good At — and What They Are Not
Model releases arrive every few weeks and marketing language makes everything sound equivalent. The practical reality is narrower, and knowing the boundary saves days.
Text to video and image to video
Text-to-video is best for exploration and mood. You describe a scene, generate several variations, and keep the one that feels right. Precision is limited because the model resolves your description in ways you did not specify.
Image-to-video is where production value lives. You supply a first frame — a rendered still, a frame grab, a photographed location, a designed graphic — and the model animates from it. Because composition, colour and framing are already locked, the output is far more controllable. Most professional work gravitates here: design the frame, then move it.
Speech, lip sync and dubbing
Several current models generate synchronised speech alongside the picture, which removes a separate dubbing step for scratch tracks and rough cuts. For a Dutch production that also needs an English version, this is a meaningful shortcut: generate the Dutch performance for the edit, then produce a dubbed variant for international distribution without booking a second voice session before the cut is locked.
Lip sync has improved dramatically for frontal, well-lit faces. It remains shaky in profile, with heavy occlusion, or with fast head movement. If dialogue is the hero of the shot, still shoot it for real.
Cleanup, upscaling and relighting
A large share of real-world AI use is unglamorous repair work: removing a boom shadow from a wide shot, extending a set wall that did not cover the frame, upscaling archive footage to match modern delivery, and relighting a pickup so it matches the master shot. This is where AI earns its keep quietly, because nobody in the audience notices and nobody in the client meeting objects.
Where generation still breaks
Hands interacting with objects, legible text inside the frame, precise physical continuity, reflections, and long unbroken takes with specific camera moves. Complex multi-character choreography falls apart quickly. The workaround is almost always the same: break the shot into shorter, simpler beats and cut around the difficult moments, exactly as you would with a practical shoot that refuses to cooperate.
Choosing a Model Without Getting Lost
There is no single best model. There is the model that fits the shot, the deadline and the budget. Build your selection around five questions.
Consistency and character continuity
If a character appears in six shots, continuity is the whole game. Options include locking a reference image and reusing it, training a lightweight adaptation on a small set of images, or generating a turnaround sheet and using stills from it as first frames for each shot. The last method is the most underrated: it costs nothing extra and gives you an art-department-style reference that also helps the client approve the look.
Duration, resolution and frame rate
Longer clips are not automatically better. Many models produce their most convincing motion in short bursts, so a five-second shot often beats a twenty-second one that drifts. For broadcast delivery you will upscale or finish at the required resolution anyway, which means the generation resolution matters less than the sharpness of the source frame you feed in.
A simple decision matrix
- Hero shots with a recognisable character: image-to-video with a locked reference, short durations, heavy take selection.
- Abstract transitions and B-roll: text-to-video, generate generously, discard ruthlessly.
- Dialogue close-ups: generate a clean plate, then composite real footage where possible.
- Set extensions and cleanups: dedicated inpainting and upscaling tools rather than full generation.
- Localisation: dubbing and lip-sync tools layered onto a locked cut.
Premium versus budget tiers
Premium tiers buy you better motion physics, cleaner textures, longer clips and faster turnaround. Budget tiers buy volume. A common mistake is using a premium model for every exploratory test, then running out of budget before the shots that matter. Explore cheaply, finish expensively.
A Production Workflow That Survives Real Deadlines
The workflow below is the one that consistently holds up when a client is reviewing, a deadline is fixed and three people need to touch the same sequence.
1. Treatment and shot list
Write the treatment as you always would. Then convert it into a shot list with a column for the intended method: live action, generated, hybrid, archive, graphic. This single column decides your budget, your schedule and your crew size. Do it before anyone opens a model.
2. Look development and storyboards
Generate still frames, not video, for the first pass. Stills are fast, cheap to iterate and easy to review in a PDF. Once the client signs off the visual direction, those approved stills become the first frames for generation. You have effectively storyboarded with production-quality images, which is a genuine change from a decade ago when storyboards were stick figures and hope.
3. Shot generation in batches
Generate in batches organised by scene, not by shot. Scene-level batching keeps lighting, wardrobe and colour drift visible and correctable. Name files with a strict convention — project, scene, shot, version — because by day three you will have hundreds of clips and no memory.
Keep a take log. A simple sheet with columns for shot, prompt or reference, model, duration, verdict and reason is worth more than any automated tagging system. The reason column is the important one: it stops you repeating the same failure twice.
4. Assembly, sound and grade
Cut generated material like any other footage. It benefits from sound design more than live action does, because generated imagery often lacks micro-texture and diegetic noise. Add room tone, footsteps, cloth movement and atmosphere. Grade generated clips to match live-action plates rather than the other way around; grading live action to match a generated look is usually a losing battle.
5. Delivery and archive
Deliver from the finished sequence, not from the generated source files. Archive the prompt or reference for every shot that made the cut, because a picky client will ask for a small change three weeks after delivery, and regenerating from a lost prompt is avoidable pain.
Working With Clients, Agencies and Broadcasters in the Netherlands
Dutch production culture is direct, well organised and allergic to vagueness. That suits AI-assisted work, provided you are equally direct about what it is.
What clients actually ask
They ask three things: does it look right, is it legally clean, and can you do it again next quarter. Answer all three in the first meeting. Show a short proof of concept, explain where the material came from, and describe how you would repeat the process at scale. Producers do not buy novelty; they buy repeatability.
Disclosure, rights and consent
Rules vary by broadcaster, platform and campaign. Establish in writing, up front, whether the client wants AI involvement disclosed, whether generated faces may resemble real people (best answer: no, unless you have explicit consent), and who owns the generated assets. Where a real performer's likeness is referenced, get a signed release covering synthetic use. Where music is used, clear it as normal — generation does not exempt you from rights.
Budget conversations
The honest framing is that AI shifts spend rather than removing it. You save on shoot days and permits; you spend on iteration time, review cycles and finishing. Present it that way and the conversation stays grounded.
Compute, Queues and Keeping the Schedule Intact
Rendering is the part nobody puts in the schedule and everybody complains about later. Generation jobs are heavy, they run in queues, and a queue that is fine at eleven in the morning is a disaster at six in the evening the day before delivery.
Practical habits that help:
- Work in parallel, not sequentially. Submit a batch, then write the next scene's shot list while it runs. Never sit and watch a progress bar.
- Generate overnight. Long or high-resolution jobs belong in the evening. Check results over morning coffee.
- Keep a local fallback. A rough cut can be assembled with placeholders so editorial and sound continue during any outage.
- Version aggressively. Save a locked cut before every experimental change.
- Watch the clock, not the counter. Track how long each scene took end to end. That number, not the tool's feature list, decides whether the method is viable for your next project.
Ten Mistakes That Sink AI-Assisted Projects
- No shot list. Generation without a plan produces attractive footage that does not cut together.
- Chasing long clips. Longer takes reveal drift. Cut more, generate shorter.
- Ignoring continuity. Wardrobe, hair and lighting changes between shots are the most common client complaint.
- Skipping stills. Approving the look in stills is faster and cheaper than approving it in motion.
- No naming convention. You will lose the one version that worked.
- Generating sound last. Poor sound design exposes weak imagery immediately.
- Using premium tiers for exploration. Explore cheap, finish expensive.
- Hiding the method. Clients forgive AI; they do not forgive surprise.
- Forgetting rights. Likeness, music and archive clearances still apply.
- No fallback plan. A clip that will not resolve must be replaceable by a practical shot or a graphic.
The Human Roles That Matter More, Not Less
AI does not flatten the crew; it redistributes the emphasis. The roles that gain weight are cinematography literacy, editorial judgement, sound design, colour, and above all direction and producing.
Someone has to decide what the film is about. Someone has to look at forty variations and say "that one, and here is why." Someone has to hold the client relationship when a scene takes longer than promised. Those are not tasks a model performs.
What changes is the ratio of craft time. Less time is spent waiting for weather and permits, more time is spent on taste, selection and finishing. A director with a strong visual instinct and a clear sense of story still outperforms a director with a shoebox of impressive clips and no throughline — dramatically so, because generation gives you infinite options, and infinite options are a trap for the undecided.
The practical implication for a Dutch studio is a change in hiring. Junior roles that were purely about gear handling increasingly overlap with prompting, asset management and version control. The best hires now are people who can argue about a frame and also organise a folder structure.
Frequently Asked Questions
Can AI-generated video pass broadcast technical review?
Yes, if you finish properly: correct resolution, frame rate, colour space and loudness, plus a delivered master assembled in an editing system rather than exported straight from a generator.
Do I still need a camera crew?
For dialogue, performance and anything requiring precise physical interaction, yes. For inserts, atmosphere, visualisations and pickups, often not.
How many takes should I generate per shot?
Plan on five to twelve for exploratory work and two to four once you have a locked reference frame. Budget your schedule around selection time, not generation time.
Is one model enough?
No. Most studios use two or three: one for exploration, one for hero shots, one for cleanup and finishing.
How do I keep a character consistent across shots?
Lock a reference image, reuse it as the first frame for every shot, keep lighting direction consistent, and avoid radical camera angles that expose the model's weaker understanding of a face.
What about audio?
Generate synchronised speech where it helps, but always finish sound in a proper audio tool with dialogue editing, foley and a loudness pass.
How do I price this work?
By deliverable and iteration rounds, not by generated seconds. You are selling finished minutes and approval cycles.
Will clients accept it?
Increasingly, yes — provided you are transparent about method and consistent about quality.
Where This Is Going
Expect three trends to shape the next stretch of work. First, tighter integration between generation and editing tools, so shots move into a timeline without manual exporting. Second, better reference control, which is the single feature that matters most to professional users: give a model a face, a wardrobe and a lighting setup, and hold it across a sequence. Third, more emphasis on finishing quality — upscaling, restoration, colour and sound — because that is where the gap between amateur and professional output is genuinely visible.
For Dutch filmmakers, the opportunity is not to replace craft with automation. It is to spend more of a fixed budget on taste and less on logistics, and to bring visual ideas to a client conversation that previously existed only in someone's head. The directors who benefit most will be the ones who treat these tools as another department — useful, limited, and worth learning properly.
Start small. Pick one scene from a project you already understand. Storyboard it in stills, generate it in short beats, cut it with real sound design, and time the whole thing honestly. That single experiment tells you more about whether AI belongs in your pipeline than any comparison table ever will.

