Why Aristotle Still Matters in an AI-Native Workflow
Generative video tools have collapsed the cost of producing images. They have not collapsed the cost of producing meaning. A model can render a rain-soaked street in seconds, but it cannot decide that the rain matters because the protagonist is about to admit she lied. That decision is still a writing decision, and the oldest useful framework for making it remains Aristotle's Poetics.
The Poetics is not a rulebook you must obey. It is a diagnostic vocabulary: plot as the arrangement of incidents, character revealed through choice, reversal and recognition as the engines of emotional payoff, catharsis as the release that makes an audience feel the time was well spent. Those ideas scale down beautifully. A three-minute short with one location and two characters carries the same structural obligations as a feature, compressed to the bone.
Structure matters more, not less, when footage is generated. Generation is probabilistic. Every shot is a negotiation with a model that has no memory of your intent, and if your structure is vague, the model will fill the gap with generic beauty: slow dolly moves, lens flares, a wistful look at the horizon, a drone shot of a city at dusk. Generic beauty does not accumulate into feeling. Structured conflict does. The practical move is to treat the script as a control layer and the generator as an execution layer, then design that control layer with the rigor a screenwriter would bring to a festival submission.
This guide walks through the whole workflow: the three-act spine for a short, mimesis translated into prompt language, hamartia and catharsis engineered on purpose, character and continuity control, prompt patterns tied to dramatic function, and the decision criteria for when AI is simply the wrong tool.
The Three-Act Spine for a Short AI Film
A short film does not need less structure than a feature. It needs structure that arrives faster, with fewer supporting characters and less exposition. The three-act spine survives because it answers three questions an audience asks unconsciously: what world am I in, what pressure is building, and what does the ending mean?
A useful ratio for a three-to-four minute short is roughly 20 percent setup, 60 percent escalation, and 20 percent resolution. On a 240-second runtime that is about 48 seconds of setup, 144 seconds of escalation, and 48 seconds of resolution. The escalation block is where almost every AI short fails, because escalation requires cause and effect, and cause and effect require you to know what changed between shot five and shot six.
Act one: establish the world, the want, and the wound in under a minute
Act one has one job: make the audience care about a specific person in a specific place with a specific problem. You do not need backstory. You need three visible facts. Where are we? What does this person want right now, in this scene? What is stopping them, and what does the obstacle cost?
Example: a night-shift parking attendant who keeps a suitcase packed under the desk. Three shots establish world (empty garage), want (she watches arrivals on a monitor, waiting for one specific car), and cost (the suitcase, visible but unexplained). No dialogue required. The suitcase is the wound.
Act two: escalate through decision, not through scenery
Escalation means the protagonist makes choices, and each choice narrows their options. This is where most generated shorts become a mood reel: ten beautiful shots of the same emotional register, none of which changes anything. Fix it with a simple test. For every shot in act two, write one sentence starting with "Because of the previous shot, she now..." If you cannot complete the sentence, the shot is decorative.
Build act two around two or three pressure turns. In a four-minute film, place them at roughly the one-minute mark, the midpoint, and the two-and-a-half-minute mark. Each turn should cost something concrete: a key, a phone number, a lie told out loud, a door closed.
Act three: compress hard and end on an image, not an explanation
Shorts rarely have room for denouement. End within fifteen seconds of your climax. The strongest ending for an AI-generated short is usually a single held image that recontextualizes the opening shot: the suitcase open, the garage lights off, the same monitor now showing nothing. Repetition with a change is the cheapest, most reliable emotional device in short-form filmmaking, and it is also extremely easy to control with generation tools because you already have the reference frames.
Mimesis: Making Generated Footage Feel Observed, Not Invented
Aristotle's concept of mimesis is often flattened into "imitation," but the useful reading for filmmakers is representation of action: art shows us people doing things and lets us infer what it means. Models are excellent at surfaces and weak at action, because action implies continuity across time. Your job is to specify behavior rather than appearance.
From abstract mood to observable behavior
Compare two prompts. Prompt A: "A lonely woman in a dim apartment, cinematic, melancholic, golden hour." Prompt B: "A woman in her forties stands at a kitchen counter, holding a phone in her left hand, thumb hovering over the screen without pressing it. Dishes in the sink behind her. Her shoulders stay level; only her thumb moves."
Prompt A produces a mood. Prompt B produces a person with a decision pending, which the audience reads as story. The second prompt is also easier to reproduce across shots, because the physical details act as anchors. Whenever a prompt contains an adjective that describes a feeling, replace it with a verb that describes a body.
Reference discipline and shot grammar
Keep a small visual bible: three to five reference images for each character, one for each location, and a written note on the light direction. Then vary only the camera. Shot grammar communicates psychology, so a character under pressure should be framed differently at the end of act two than at the start. Wide shots isolate, close-ups implicate, and static frames feel like observation while handheld frames feel like anxiety. Decide this in the script stage so generation does not decide it for you.
Hamartia, Reversal, Catharsis: Engineering Emotional Payoff
Hamartia is usually translated as a fatal flaw, and the cinematic version is narrower: a character's own strength, pushed too far, becomes the thing that destroys them. A loyal sister who cannot stop covering for her brother. A careful archivist who refuses to destroy anything, including evidence that should burn. Flaws that are also virtues produce the most interesting shorts, because the audience understands the choice.
Give the protagonist a flaw that causes the plot
Test your premise with one question: does the protagonist's flaw cause the central problem? If the problem is caused by an external villain or random accident, the story becomes a sequence of events rather than a character study. Rewrite until the flaw is load-bearing. This is especially important in AI shorts, because a passive protagonist plus gorgeous generated imagery reads as a screensaver.
Build the reversal and recognition beat explicitly
Reversal is a change in fortune; recognition is a change in knowledge. The strongest short endings deliver both in the same beat: the character learns something and, because of that knowledge, loses or gains something irreversible. Write that beat as a single line in your beat sheet, for example: "She finally opens the suitcase and finds her own handwriting on the note inside." Then build the shot list backward from it.
Catharsis, the release, is a pacing problem more than a writing problem. It needs silence. In an AI short, plan the last six to ten seconds as one shot with no camera movement, minimal score, and a held frame. Audiences forgive a lot if the final image lands.
A Practical Pipeline: From Logline to Locked Cut
Step 1: logline and dramatic question
Write one sentence: a character, a want, an obstacle, and a cost. Then write the dramatic question an audience should be asking at the end of act one. If you cannot state it in nine words, the short does not have a spine yet.
Step 2: beat sheet into a shot list
Convert each beat into one to three shots, aiming for a total of forty to seventy shots for a four-minute film. Number them and tag each with its dramatic job: setup, pressure, turn, reveal, release. This tag is the most valuable column in your document, because it lets you cut footage that is beautiful but structurally unemployed.
Step 3: keyframes and character reference sets
Before generating any moving footage, produce stills. Lock the character's face, wardrobe, and key props as images. A single consistent still can seed dozens of shots, and it is far cheaper to reject a still than a full clip. Approve the visual bible in this phase, then stop redesigning.
Step 4: generation passes by dramatic function
Generate in passes grouped by function, not by scene order. All setup shots together, then all pressure shots. This keeps your prompt vocabulary stable while you work, and it surfaces continuity problems early. Keep two takes per shot minimum: one literal, one looser. The looser take frequently gives you the gesture that makes the cut.
Step 5: assembly, temp sound, and the first real cut
Cut picture with temp music and scratch ambience. Watch it once without pausing and mark every moment your attention drifts. Those moments are almost always structural, not visual. You will usually fix a drifting act two by deleting shots rather than generating new ones.
Continuity Control: Faces, Wardrobe, Light, and Space
Continuity is the technical tax on AI filmmaking. Face drift, changing jacket colors, and light that flips direction between shots are the three most common tells. Solve them with discipline rather than hope.
- Identity: keep one approved portrait per character and reuse it as a reference in every shot prompt. Describe the face using three permanent traits and never change the wording.
- Wardrobe: describe clothing in a fixed phrase you copy verbatim, including fabric and color. Do not paraphrase it between prompts.
- Light: state direction, quality, and color temperature in every prompt. "Light from screen-left, soft, cool ambient" is more useful than "moody lighting."
- Space: define an axis for each location and keep the camera on one side of it, the same principle as a stage line in classical continuity editing.
Keep a continuity log with one row per shot: character ID, wardrobe phrase, light phrase, location ID, and dramatic job. It feels bureaucratic for a four-minute film and it will save you an entire regeneration cycle.
Prompt Patterns That Encode Dramatic Structure
The most underused idea in AI filmmaking is that a prompt can carry dramatic function, not just visual description. Structure your prompts in four slots: subject and action, framing, light, and continuity anchors. Then vary only what the dramatic job requires.
| Dramatic job | Prompt emphasis | Shot choice |
|---|---|---|
| Setup | Environment detail, stillness | Wide or medium, static |
| Pressure | Small physical action, hands | Medium close, slow push |
| Turn | A single decisive movement | Close-up, short lens |
| Reveal | Object or face, minimal motion | Insert or extreme close-up |
| Release | Negative space, held frame | Wide, locked off |
Two habits make this work. First, one action per shot: models handle a single clear verb far better than a chain of them. Second, negative prompts for tone: exclude "smiling," "crowd," or "bright daylight" when the scene requires isolation, and the output will read as intentional rather than accidental.
Common Mistakes and Their Fixes
Mistake: writing for visuals instead of for change. If every shot is a different angle of the same emotional state, the film stalls. Fix: require each shot to change either information or power.
Mistake: too many characters. Shorts with five speaking roles in three minutes have no room for character. Fix: reduce to two, and let the third exist only as a voice or a photograph.
Mistake: exposition in dialogue. AI-generated dialogue often reads flat. Fix: move information into props, blocking, and the order of shots. Let the audience assemble the facts.
Mistake: inconsistent pacing. Generated clips arrive in whatever length the tool prefers. Fix: cut to rhythm in the edit and use a temp score to enforce it.
Mistake: solving structural problems with more generation. If act two drags, generate nothing for a day and cut instead. Deleting the three weakest shots is usually the whole fix.
Mistake: an ending that explains. Narration over the final shot is almost always a sign that the image did not do its job. Fix: rewrite the final shot, not the narration.
Sound and Cutting Rhythm: Where Emotion Lives
Picture gets the attention; sound gets the feeling. For a short, plan three layers: a bed (room tone, rain, traffic), a signal (the sound the protagonist is listening for), and a score that enters late and leaves early. The signal layer is a structural device, not decoration: if a phone buzzes at 0:40, its absence at 3:20 does narrative work.
Cutting rhythm should follow act structure. Setup cuts can breathe at three to five seconds per shot. Escalation should shorten to two to three seconds, with one deliberate long take just before the turn to create contrast. The release should be the longest shot in the film. If you edit in a tool that supports markers, place one at every pressure turn so you can see the rhythm on the timeline rather than feel it vaguely.
Decision Criteria: AI, Live Action, or Hybrid
Use generated footage when the scene depends on imagery that would be expensive, impossible, or ethically complicated to shoot: dream states, historical settings, surreal interiors, crowd scale, or a single actor playing multiple roles. Use live action when the scene depends on a performance beat so fine that a fraction of a second of hesitation carries the meaning. Physical comedy, overlapping dialogue, and improvised conflict remain hard for generators because they require precise timing across bodies.
Hybrid is usually the best answer for a short. Shoot the emotional close-ups with a real camera and a real actor, even on a phone, then generate the world around them. The generated footage provides scale and texture; the live footage provides the truth. Audiences forgive a mismatch in image quality far more readily than they forgive a mismatch in feeling.
Frequently Asked Questions
How long should an AI-generated short be? Two to five minutes is the sweet spot. Under ninety seconds it is difficult to establish character and deliver a turn; over six minutes the production cost grows faster than the audience's patience unless the story has real complexity.
Do I need a full screenplay? No. A one-page beat sheet plus a numbered shot list with dramatic function tags is enough for most shorts. What you cannot skip is the ending beat, written as a single concrete line before you generate anything.
How many shots should I generate per finished shot? Budget three to five attempts for complex action, one to two for static inserts. Track wins and failures so you learn which shot types your chosen model handles reliably.
Can I fix a weak story in the edit? You can improve pacing, tighten transitions, and remove dead shots. You cannot invent a reversal that was never generated. Structure is decided before production.
What if my character's face drifts between shots? Return to a locked reference image, keep the descriptive phrasing identical, and reduce the amount of action in the prompt. Complexity is the main cause of identity drift.
How do I know when the short is finished? When you can watch it once without wanting to explain anything to the viewer. If you feel the need to add context, the story is still missing a beat, and the fix is a shot, not a caption.
Is it worth writing a logline for a two-minute film? Yes. A logline is a test: if the sentence is boring, the film will be too. Rewriting one sentence costs minutes; regenerating a short costs days.
How much should the audience be told? Less than you think. Give them the want and the obstacle, and withhold the reason. Recognition lands hardest when the audience arrives at the same conclusion a beat before the character does.

