Short films built for TikTok obey different rules
A short film made for TikTok is not a festival short that happens to be cropped vertical. It is a different object with different physics. The viewer is holding a phone, watching with the sound on but the attention elsewhere, and one thumb-flick away from leaving. Everything about the production process has to serve that reality: the first 1.5 seconds, the vertical frame, the audio-first design, the loop at the end.
Generative video models make this kind of production genuinely practical. A single creator can now produce a 30-second narrative piece with a consistent character, three distinct locations, and a soundtrack the same afternoon. The hard part is no longer access to motion images. The hard part is directing a stack of models so that the result looks like one film instead of a collection of unrelated clips.
That is what this guide covers: a repeatable workflow for scripting, generating, and finishing shareable vertical short films using multiple AI models, plus the decision criteria that keep you from wasting an afternoon on the wrong tool.
Start with the format, not the tool
The most common failure mode is opening a generation interface before deciding what kind of film you are making. Models reward clear constraints. Formats provide them.
These five formats cover most of what performs well in vertical feeds:
- Single-idea explainer (20–35s). One concept, one narrator, six to eight shots. Ideal when you want a hook that is entirely verbal and visuals are supporting evidence.
- Character vignette (30–45s). A small emotional beat with one character. Best showcase for consistency controls because the audience stares at one face for the whole runtime.
- Visual punchline (12–20s). A setup and a twist. Very short shot count, high replay rate, forgiving of imperfect detail.
- Serialized micro-episode (45–60s). Part one of a series. The cliffhanger replaces the ending and drives follows rather than just views.
- Product or process demo (20–30s). Macro shots, hands, materials, motion. The strongest format for pure text-to-video because it rarely requires a face.
Match runtime to shot count before you write
A useful planning rule: budget roughly 3–5 seconds per generated shot, then add cutaways. A 30-second film is not 10 shots of 3 seconds; it is usually 7–9 generated shots plus 4–6 quick insert or text cards.
| Runtime | Generated shots | Insert/text cards | Typical cut pace |
|---|---|---|---|
| 15s | 4–5 | 2–3 | every 1.3s |
| 25s | 6–7 | 3–5 | every 1.8s |
| 40s | 9–11 | 5–7 | every 2.2s |
| 60s | 13–16 | 7–10 | every 2.5s |
Locking this table before generation is what keeps a project finite. Without it, "one more shot" becomes the default and the film never ships.
Build the pre-production package first
Generating before pre-production is expensive in both time and render budget. Three artifacts prevent almost every problem downstream.
The one-page script
Write the film as a page with five columns: timecode, voiceover line, on-screen text, shot description, and shot intent (why this shot exists). If a shot has no intent beyond looking good, cut it. Vertical films die from decorative footage.
A 25-second character vignette might read like this:
- 0:00–0:03 — VO: "I moved to this city for the quiet." Shot: wide rooftop at blue hour, small figure standing at the edge. Intent: establish loneliness and scale.
- 0:03–0:07 — VO: "Turns out quiet is just noise you haven't met yet." Shot: medium close-up, character turning as traffic light flares. Intent: introduce the turn.
- 0:07–0:13 — no VO, ambient. Shot: slow push-in on the character's face, city bokeh behind. Intent: hold the emotion without words.
- 0:13–0:19 — VO: "So I started listening to it." Shot: low angle, character walking into a crowd, camera tracking. Intent: release tension.
- 0:19–0:25 — VO: "Now the city is loud and I sleep fine." Shot: same rooftop framing as the opening, daytime. Intent: visual rhyme to close the loop.
The character and environment bible
Write down, in text, what never changes: hair length and color, jacket, a scar, the exact time of day, the color temperature of the key light. This text becomes part of every prompt. Vague bibles produce drifting characters, and drifting characters are the single most visible sign of an AI-made film.
The shot ledger
A spreadsheet with one row per shot: shot ID, scene, model used, prompt, seed, reference images, take number, status, and notes. It sounds bureaucratic. It saves you from regenerating a shot you already perfected three days ago, and it makes it possible to hand the project to an editor.
Choose the right model for each shot, not one model for the film
Multi-model production is not about collecting tools. It is about assigning each shot to the model whose strengths match the shot's demands. Four criteria do most of the work:
- Motion complexity. Simple camera moves and walking characters are handled well by nearly every model. Complex physics, crowds, water, and fight choreography separate the field quickly.
- Consistency controls. Seeds, reference images, motion brushes, character training, and last-frame continuation matter more than raw output quality for narrative work.
- Duration per generation. Models with short maximum clip lengths push you toward more cuts, which is often fine for TikTok but bad for an uninterrupted emotional beat.
- Text and hands. On-screen typography generated inside a model is usually unusable. Plan to add all text in the editor, and avoid shots that require readable signage.
A practical assignment matrix
| Shot type | Best approach |
|---|---|
| Establishing environment, no people | Text-to-video |
| Character close-up with emotional beat | Image-to-video from a locked still |
| Consistent character across five shots | Generate one hero still, then image-to-video with that still as reference |
| Camera move matching a live reference | Video-to-video with motion guidance |
| Dialogue or narration synced to a face | Dedicated lip-sync tool layered over a generated clip |
| Product macro, texture, liquid | Text-to-video, short clips, heavy cutting |
| Crowds, sports, complex physics | Highest-tier model available, expect multiple takes |
The one-model trap
Using a single model for everything feels simpler and usually costs more time. You end up compromising the close-ups to accommodate the crowd shots. The alternative is a small, stable stack: one workhorse for environments, one for character shots, one for animation and lip-sync, one upscaler, one editor. Four tools, learned well, beat twelve tools used casually.
Continuity is a production system, not a prompt trick
Continuity is where AI short films are won. Audiences forgive soft detail and odd physics; they do not forgive a jacket that changes color between shots.
Lock the still before you animate
For any character-driven shot, generate stills first. Stills are fast and cheap to iterate on. When a still is right — face, wardrobe, lighting — save it, name it, and treat it as a locked asset. Only then run image-to-video. Animating an unapproved still is the most common source of wasted render budget.
Reuse frames deliberately
The last frame of shot 4 can be the first frame of shot 5 when the camera is meant to continue. This is the cheapest continuity tool available: no prompt engineering required, and the transition becomes invisible instead of a hard cut.
Keep a lighting and lens rule
Pick one color temperature per scene and one lens character (wide, normal, long) and write it into every prompt for that scene. Scene-level grading in the editor then only needs a single adjustment layer instead of shot-by-shot correction. Films that look "off" usually have three different key-light directions across six shots.
Write a continuity checklist per scene
Before generating a scene's shots, list: character appearance, wardrobe, time of day, weather, lens, camera height, and color palette. Run the list against each prompt. It takes ninety seconds and prevents reshoots.
Generate in batches and keep a take discipline
Random generation order produces random results. Batch by scene and by location so that when you tune a prompt, the tuning benefits the whole scene.
A workable batch process:
- Generate stills for the entire scene.
- Select the two best stills per shot.
- Animate both contenders at low resolution.
- Pick the winner and re-render at full resolution.
- Upscale only approved shots.
This low-resolution-first loop typically cuts total generation time in half. It also changes your relationship with takes: instead of chasing perfection across twenty attempts, you accept that take two of a low-res pass is the one, and you move on.
Name files so future-you can find them
Use a consistent convention: project_scene_shot_take_model. Sort by shot ID and you have an instant timeline of every asset you own. This is the difference between a project you can revise next week and a folder of mystery clips.
The edit is where the film actually appears
AI generation produces material. Editing produces the film. The post-production pass deserves at least as much attention as the prompts.
Vertical framing and safe zones
Work in 1080×1920. Keep faces in the upper-middle third, keep critical text at least 12% away from top, bottom, and right edges where platform UI sits. If a clip was generated widescreen, do not simply crop it — reframe with intentional movement, or regenerate vertical, because center-cropped widescreen usually cuts the subject in half.
Sound design carries more weight than visuals
Vertical viewers tolerate average imagery with great sound far more readily than the reverse. Layer three things: a bed (music or ambient), punctuation (transitions, whooshes, impacts), and a voice. Normalize dialogue to a consistent level, keep music well under the voice, and check the mix on a phone speaker, which is where nearly everyone will hear it.
Captions and pacing
Burned-in captions raise completion rates for most vertical content. Keep them to two lines, high contrast, and short phrases. Cut the first three seconds faster than the rest — roughly one cut every 1.2–1.5 seconds — then settle into a slower rhythm. End on a frame that loops cleanly into the opening shot if the format supports it.
A full production day, start to finish
Here is how a single 30-second film comes together in one working session.
Hour 1 — Script and plan. Choose a format, write the one-page script, build the shot ledger with six to nine rows.
Hour 2 — Stills. Generate and lock character and environment stills. Approve wardrobe, lighting, and lens. Save everything with names you will recognize later.
Hour 3 — Animation. Batch-generate low-resolution clips per scene, then re-render the winners at full resolution.
Hour 4 — Assembly. Import into the editor in shot order, trim to the timing plan, add camera-consistent transitions, and lay a temp music bed.
Hour 5 — Sound and captions. Record or generate voiceover, sync lip movement where needed, add punctuation effects, write captions.
Hour 6 — Grade, QC, export. Apply one grade per scene, run the checklist, export at 1080×1920, 30fps, high bitrate.
Two films a week at this pace is sustainable for one person. Four is possible if you reuse environments and characters from an existing series.
Quality control checklist before you publish
Run this every time. It takes four minutes and prevents the majority of embarrassing uploads.
- First frame contains a reason to keep watching — motion, a face, or a question.
- No shot is longer than five seconds unless it is deliberately a held beat.
- Character appearance identical between shots in the same scene.
- No readable generated text anywhere in frame.
- Hands and faces checked frame-by-frame in fast shots.
- Audio peaks controlled; voice intelligible on a phone speaker.
- Captions accurate, timed, and inside safe zones.
- No unintended brand marks or watermarks.
- Ending loops or lands a clean final beat.
- File exported vertical with no letterboxing.
- Cover frame chosen deliberately, not defaulted to second one.
Common mistakes and how to fix them
Generating before scripting. Fix: never open a generation tool until the one-page script exists.
Mixing visual styles. Photoreal in shot one, stylized in shot four. Fix: declare one visual language per film and write it into every prompt.
Over-relying on long shots. Fix: cut longer, or split a long beat into two angles.
Ignoring the first second. Fix: write the hook before the story, then build the story to justify it.
Animating unapproved stills. Fix: stills first, animation second, always.
Forgetting audio until the end. Fix: temp music and a scratch voiceover during assembly, not after picture lock.
Chasing a perfect single shot. Fix: set a take limit of three per shot and move on. A finished film outperforms an unfinished masterpiece every time.
Keeping time and render budget under control
Costs in AI video come from three places: generation volume, resolution, and re-work. You control all three.
- Iterate at low resolution. Preview passes at reduced settings are dramatically cheaper per attempt and reveal 90% of problems.
- Approve in stills. Stills are the cheapest iteration surface you have.
- Set a take cap. Three takes per shot, then the ledger records the best available option.
- Reuse assets. Build a personal library of environments, props, and character stills. Series content should be 30–40% reused material.
- Upscale once. Only approved shots go through enhancement.
If you plan a monthly output of eight films, estimate the render volume for a single film first and multiply by eight before committing to a publishing schedule. Most creators overestimate what they can produce and underestimate how much time the edit consumes.
Frequently asked questions
How long should an AI short film be on TikTok?
For narrative work, 25–45 seconds is the sweet spot: long enough for one emotional turn, short enough to keep completion rates high. Pure punchline formats work at 12–20 seconds. If your story needs more than 60 seconds, serialize it rather than extending a single upload.
Can AI video keep a character consistent across many shots?
Yes, but consistency comes from process, not from any single model. Generate a locked hero still, use it as a reference for every shot in the scene, repeat the wardrobe and lighting description word-for-word in every prompt, and reuse the last frame of a clip as the first frame of the next when the camera is meant to continue.
Do I need several paid tools?
Usually two or three paid tools and one free editor is enough: a workhorse video model, a character or image-to-video model, a lip-sync tool if you have dialogue, and a capable editor with good caption tools. Adding more tools rarely improves output — better prompts and stricter selection improve output.
Should I generate dialogue audio in the model?
No. Generate picture first, then record or synthesize clean voiceover separately and sync it in the editor. Generated audio inside video models is still inconsistent and hard to mix. Clean post-production audio is the fastest way to make AI footage feel professional.
What about music?
Use tracks you have clear rights to, whether from a licensed library or a royalty-free source that permits commercial social use. Keep a record of the license for each track. Muted uploads undo weeks of work in a single strike.
How many takes should I allow per shot?
Three at full resolution. Unlimited at preview resolution, but only during a single focused tuning session — otherwise you are polishing a shot that will be on screen for 1.8 seconds.
Can I build a series with the same character?
Yes, and it is the highest-leverage thing you can do. Save your character stills, environment plates, and prompt blocks in a reusable project file. Series content compounds: the audience recognizes the character, and your production time per episode drops sharply after the third film.
The core discipline is simple. Decide the format, write the script, lock the stills, assign each shot to the right model, animate in batches, and treat the edit as the place where the film is actually made. Models will keep improving. A workflow that produces finished, shareable films every week will keep working regardless of which model is fastest this month.

