Limited Time Offer: Get 50% OFF your first month of Pro & Ultra plans 🎉

Turn Long YouTube Videos Into Short Vertical Clips With AI

Sep 12, 2026

Why turning wide video into vertical shorts is its own craft

Long-form video and vertical short-form video may share the same footage, but they share almost nothing else. A twenty-minute upload earns attention through accumulation: a thesis, supporting examples, and a slow build toward a payoff. A vertical short earns attention through compression: one idea, one emotional turn, and a reason to rewatch, delivered before a thumb moves.

That difference explains why "cut the best part and post it" so often underperforms. The best part of a long video is frequently the part that depends most on context. Strip out the setup and the clip stops making sense, so viewers leave in the first two seconds — and the algorithm reads that exit as a weak signal.

A workable repurposing workflow does three jobs at once. It finds moments that stand alone. It rebuilds the visual framing for a phone held upright. And it adds enough structure — captions, pacing, a hook — that a cold viewer understands the clip with zero prior knowledge. AI helps at every one of those steps, but only if a human keeps control of the question that matters most: which moments are worth posting.

The five-stage pipeline at a glance

Before comparing tools, sketch the pipeline. Most creators who struggle with repurposing are skipping a stage rather than using bad software.

  1. Audit. Skim the source with a transcript open and mark where the argument turns.
  2. Select. Score candidate moments against a rubric, and pick more than you need.
  3. Reframe. Convert the wide frame into a vertical composition that keeps the subject readable.
  4. Polish. Captions, audio cleanup, hook text, and an ending that loops.
  5. Sequence and test. Publish in batches, compare retention, and feed what you learn back into stage two.

A dense, conversational 20–30 minute video typically yields 8–15 usable shorts. Interview and podcast footage yields more, because every answer is already a self-contained unit. Tightly scripted tutorials yield fewer, because each segment leans harder on the material before it.

Expect 45–90 minutes per finished short while you are learning the workflow, dropping to 10–20 minutes once you have caption presets, reframing habits, and a template you trust. The goal is not speed for its own sake; it is speed that leaves you room to be choosy about which clips ship.

Stage one: audit the source before you cut anything

Read the transcript first, not the timeline

Open the transcript and read it end to end. Reading is faster than scrubbing, and it forces you to notice structure: where a claim is introduced, where it is challenged, where a story lands. Highlight anything that produces a reaction in you — surprise, disagreement, a laugh, a clear before-and-after.

Then go back to the timeline for those timestamps only. You are checking whether the moment has visual life: a facial expression, a gesture, a screen demonstration, a prop. A brilliant sentence delivered over a static talking head with no movement is harder to sell than an average sentence with a visible reaction.

Score candidates with a simple rubric

Give every candidate a quick 1–5 score on five dimensions:

  • Standalone clarity. Does it make sense with no setup?
  • Emotional charge. Does it contain tension, surprise, or a strong opinion?
  • Specificity. Are there numbers, names, or concrete examples?
  • Visual interest. Is there something to look at besides a face?
  • Duration fit. Can it land in 20–60 seconds?

Anything scoring below 3 on clarity or charge is usually not worth the edit, no matter how good it sounded in context. Keep the scoring sheet. It becomes your editorial memory, and it stops you from re-litigating the same choices every week.

Build a clip inventory instead of editing one at a time

Collect candidates in a simple table: timestamp range, working title, one-line summary, score, and status. Twenty rows take about twenty minutes to build and give you a month of publishing material. Editing from an inventory also keeps you from over-investing in a single clip that turns out to be undeliverable.

Stage two: let AI surface highlight candidates

Transcript-first selection beats frame-first selection

AI highlight detection works best when it reads language, not pixels. A transcript-plus-topic model can identify self-contained arguments, question-and-answer pairs, list-style explanations, and emotional peaks far more reliably than a model watching raw frames.

So the practical order is: transcribe, segment by topic, then score. Many editors now transcribe automatically and let a model propose topical boundaries. That output is a starting map, not a final cut list.

Prompt patterns that produce usable candidates

When you hand a transcript to a language model, be explicit about the constraints you care about. Vague requests like "find the best parts" produce vague lists. Better:

  • "Identify passages between 30 and 60 seconds spoken at a natural pace that make a complete point without referring to earlier context."
  • "Flag every moment where a specific number, tool name, or concrete example appears."
  • "Group the transcript into topics and tell me which topics could work as a standalone short for someone who has never seen this channel."
  • "For each suggestion, quote the opening sentence and explain in one line why it stands alone."

That last instruction matters. When the model has to justify a pick, weak suggestions become obvious, and you can discard them in seconds.

Keep the veto, always

AI is good at finding plausible moments and bad at knowing your audience. A clip that looks great on paper may be off-brand, redundant with something you posted last month, or attached to a claim you would rather not amplify out of context.

Treat model output as a shortlist. Your job is the veto. A useful habit: for each candidate, ask whether a stranger would understand it, feel something, and know what to do next. If any of the three is missing, rewrite the framing or drop the clip.

Stage three: reframing wide footage into a vertical frame

Auto-reframe versus manual keyframes

Automatic reframing tracks the subject and crops the frame, and it has improved a great deal for single-speaker footage. It struggles when two people trade lines, when the subject moves suddenly, or when on-screen text sits outside the crop.

A workable rule: use auto-reframe as a first pass, then manually correct any clip where the subject leaves the frame, where the mouth is cropped, or where a gesture is cut off. Budget a keyframe or two per clip for talking heads and more for anything with movement.

Screen recordings, slides, and demo footage

For tutorials, the wide frame is often the point — a wide shot of a whole interface is unreadable on a phone. Three practical fixes:

  • Punch in. Crop to the region of the screen where the action is happening, then add a subtle zoom on the moment that matters.
  • Split the frame. Put the speaker in the upper third and the screen capture below, so the viewer gets a face and the detail at the same time.
  • Re-shoot the beat. If a step genuinely cannot be shown vertically, re-record a five-second vertical close-up of just that step. It is usually faster than fighting the original frame.

Respect platform interface safe zones

Whatever you export, keep essential text out of the bottom band where captions, usernames, and buttons appear, and out of the top band where platform headers sit. A useful habit is to build a template with semi-transparent guides on the top and bottom, then place hook text inside the safe area.

Stage four: captions, audio, and the first three seconds

Caption styling that survives muted autoplay

Most short-form viewing starts silent, so captions are not decoration. Treat them as the primary script. Keep them short — one to three words per line for punchy content, up to five for explanatory content — and place them near the vertical centre so the eye does not travel.

Choose one font, one weight, and one highlight colour for emphasis, and reuse them. Consistency makes your clips recognisable while scrolling, which is worth more than novelty.

Clean up speech with a gentle hand

AI noise removal and level matching are worth using, but aggressive processing makes voices sound thin and robotic. Start with the lightest preset that removes room rumble and hiss, then normalise so quiet speakers do not vanish under music. If you add a music bed, keep it well under the voice — around 15–20 percent of perceived loudness is usually enough.

Hook patterns that fit the content

A hook is not a slogan; it is a promise about what happens in the next thirty seconds. Patterns that work across niches:

  • Contradiction. "Everyone says to post more. That is the problem."
  • Numbered promise. "Three settings that fixed my export quality."
  • Mid-action open. Start on the moment of tension, then rewind one sentence to explain.
  • Cost of ignoring. "This one mistake cost me a week of edits."

Whichever pattern you use, avoid a five-second intro. The clip should already be moving when it starts, and the payoff should arrive before the halfway mark so viewers have a reason to finish.

Stage five: sequencing, cadence, and testing

Batch production beats daily panic

Produce in batches of five to ten shorts. Editing in one focused session keeps your caption style, framing, and pacing consistent, and it means a single weak day does not break your posting rhythm.

Schedule with a buffer of at least a week so you can respond to performance without improvising under pressure.

What to test first

Change one variable at a time, and give each test enough clips to mean something:

  1. Hook type. Contradiction versus numbered promise.
  2. Length. 20–30 seconds versus 45–60 seconds.
  3. Caption style. Minimal lower-third versus large centred captions.
  4. Opening frame. Face-forward versus action shot.

Read the retention curve, not the view count

Views tell you distribution; retention tells you craft. Look at the first three seconds, the midpoint, and the final second. A steep early drop suggests a weak hook or a slow first frame. A drop at the midpoint usually means the clip needed a tighter edit or a second beat of tension. Viewers leaving right at the end is normal and often healthy — a rewatch loop can offset it.

A pre-export quality checklist

Run every clip through the same short list:

  • The first frame shows something happening, not a title card.
  • The opening line makes sense with no context.
  • Captions stay inside the safe zone and never cover a face.
  • Audio is normalised and free of clipping.
  • Any on-screen text is legible on a small phone screen.
  • The clip has one idea, not three.
  • The ending either lands a punchline or sets up a natural rewatch.
  • The filename and caption include the topic keywords you want to be findable for.

Print it, pin it, or paste it into your project notes. Checklists save more time than shortcuts.

Mistakes that quietly kill retention

Cutting at sentence boundaries only. Silence is not the enemy; dead air is. Trim breath pauses and filler words, but leave enough room for a joke or a reaction to land.

Over-tightening. Clips that jump every 400 milliseconds feel frantic and hide the speaker's personality. Let a shot breathe for two or three seconds when the content is emotional.

Ignoring the source's weak audio. If the original recording is bad, no amount of processing will fully fix it. Re-record the audio for that clip or choose a different moment.

Posting near-duplicates. Two clips from the same minute cannibalise each other. Space related clips across weeks.

Losing the thread of the original. A clip that distorts the source's meaning creates cleanup work later. Stay honest about context, and link the full conversation in the description when the topic invites disagreement.

Choosing tools without overbuying

A repurposing stack only needs to cover four capabilities, and each has several reasonable options.

Transcription and highlight detection. Any accurate speech-to-text engine plus a language model gets you most of the way. Test accuracy on five minutes of your own audio, especially if you have accents, jargon, or multiple speakers.

Vertical reframing. Look for subject tracking, manual keyframe correction, and the ability to export a reusable template. If you cannot correct a bad auto-crop in under a minute, the tool will slow you down.

Captioning. Prioritise editing speed and style consistency over exotic animation. You will adjust caption text far more often than you will change fonts.

Audio repair. A light denoiser, a loudness normaliser, and a music ducking feature cover nearly everything. Anything more elaborate is usually a distraction from editing decisions.

Try each candidate on one real clip before committing. One clip reveals more about a tool's friction than a week of demos.

Frequently asked questions

How many shorts should one long video produce?

Aim for 8–15 candidates and expect to publish 5–10. Quality control matters more than volume; a weak clip dilutes the perceived quality of everything around it.

Is automatic highlight detection good enough to skip the audit?

No. It is a useful first pass that saves you scrubbing, but it misses tone, brand fit, and the context your audience already has. Combine it with a manual skim of the transcript.

Do I need to shoot vertically, or can I crop?

Both work. Cropping is faster and preserves the original performance. Shooting some footage vertically — especially cutaways and screen close-ups — gives you flexibility and better resolution for the moments that matter most.

How long should a vertical short be?

Most successful clips land between 20 and 60 seconds. If a moment needs 90 seconds, split it into two clips with a clear handoff, or tighten the explanation and cut the rest.

What if my source video is a screen recording with no talking head?

Punch in on the interface, add large captions that narrate the step, and use short zooms to guide the eye. Re-record a vertical close-up for any step that cannot be read in a narrow crop.

Should captions match the spoken words exactly?

Mostly yes, with light cleanup for filler words and false starts. Lightly condensed captions read faster and still feel accurate.

Making the workflow stick

The real advantage of an AI-assisted repurposing workflow is not that it removes editing. It is that it moves the expensive part of your attention to the decisions that actually determine performance: which moment to publish, how to open it, and whether a stranger understands it in three seconds. Automate transcription, reframing, and captions — then spend the time you save on being selective. A single strong clip that fits its audience will always outperform ten clips produced on autopilot.

Alexander

Alexander