Why Instagram video specifications decide your reach
Every clip you publish passes through a pipeline you do not control. Instagram decodes your upload, rescales it, and re-encodes it for several distinct surfaces: a full-screen vertical player, a cropped feed preview, a muted autoplay tile, a compressed preview inside a message thread. When the source file fights that pipeline — wrong aspect ratio, soft master, brittle audio — the finished result looks amateurish on exactly the screens where attention is hardest to earn.
That makes specification work part of creative direction rather than an afterthought for editors. A carefully lit product shot with burn-in captions near the bottom edge can end up buried under interface elements. A crisp action sequence uploaded with an aggressive bitrate ceiling turns into a smear of blocks in the first second. Viewers rarely articulate what is wrong; they simply scroll.
The upside is that the rules are learnable and stable. Once you know the canvases, the resolution targets, the frame rate conventions, and the compression behaviour, you can build a small set of export presets and stop guessing. What follows is a practical tour of the numbers, then a repeatable production workflow — including where generative video tools genuinely help — so your output looks intentional on every surface.
The Instagram placements and their canvases
Reels
Reels are the vertical full-screen format and the one most creators care about. Design at 1080 × 1920 pixels, 9:16. The player covers the whole screen, so your composition has to work edge to edge, but bands of that canvas are effectively reserved by the interface: the handle, caption, and audio strip at the bottom, and the navigation bar at the top. Treat the outer frame as decoration and keep meaning in the middle.
Duration is flexible, but the practical sweet spot for retention is short. A tight 15–30 second piece usually outperforms a sprawling two-minute one unless the story genuinely needs the length. Loops help, so ending on a frame that flows back into the opening frame encourages replays.
Stories
Stories share the 1080 × 1920 canvas but a different psychology. They are watched quickly, often with a thumb hovering, and they expire. The bottom third is heavily obscured by the reply bar and stickers, and the top is crowded by the progress bar and profile header. Text should sit in the upper-middle band, large enough to read on a phone held at arm's length.
Because Stories are consumed rapidly, pacing matters more than polish. A change of visual state every 1.5–2 seconds holds attention. Avoid long static shots with a voiceover; use motion, cuts, or animated text instead.
Feed video and carousel video
Feed video is where 4:5 (1080 × 1350) wins. It occupies more vertical space than a square post while remaining robust to cropping in different navigation contexts. Square (1080 × 1080) is still safe and useful for grid consistency. If you plan to place a video into a carousel, match the aspect ratio of the other slides so the swipe does not jolt.
A common mistake is uploading a 16:9 horizontal clip into the feed and letting the platform letterbox it. Letterboxing wastes vertical real estate and reads as repurposed television content. If you must use horizontal footage, add a designed background, a blurred fill, or split-screen framing rather than leaving black bars.
Long-form and live
Longer horizontal video and live broadcasts still favour 16:9 at 1920 × 1080. Live streams should be encoded conservatively, because a marginal connection plus a high bitrate produces freezes; 1080p at 30 fps with a moderate bitrate is more reliable than 4K at 60 fps.
| Placement | Aspect ratio | Recommended resolution | Notes |
|---|---|---|---|
| Reels | 9:16 | 1080 × 1920 | Bottom ~20% and top ~10% covered by UI |
| Stories | 9:16 | 1080 × 1920 | Heavy bottom overlay; fast pacing |
| Feed video | 4:5 | 1080 × 1350 | Best vertical footprint without cropping |
| Square post | 1:1 | 1080 × 1080 | Grid-friendly |
| Horizontal / live | 16:9 | 1920 × 1080 | Conservative bitrate for stability |
Aspect ratio, resolution, and frame rate: the numbers that matter
Choose the aspect ratio before you shoot
Aspect ratio is the decision that costs the most to fix later. Cropping 9:16 footage into 4:5 loses the top and bottom of the frame; expanding 4:5 into 9:16 means inventing pixels. Decide the destination first, then frame for it. When one shoot must serve several placements, compose with a protected rectangle in the middle of the frame that contains all essential information, and treat the surrounding area as flexible margin.
Resolution: master higher than you publish
Publish at 1080 pixels on the short edge, but master higher when you can. A 2160 × 3840 vertical master gives you headroom for stabilisation, punch-ins, reframing, and future surfaces. Downscaling a clean 4K master to 1080p preserves detail and reduces aliasing; upscaling a soft 720p master invents mush.
Uploading an extremely high resolution does not guarantee a better result. Very large files can be re-encoded more aggressively, and heavy noise or fine texture is expensive to compress. A clean 1080p file sometimes looks better than a noisy 4K one.
Frame rate: pick a convention and stay in it
Thirty frames per second is the safe default for social video. It is smooth enough for motion, standard on phones, and matches most platform pipelines. Twenty-four frames per second gives a filmic feel and works well for narrative or lifestyle content, but fast camera movement can strobe. Sixty frames per second suits sport, product spins, and gaming, and it survives slow-motion reinterpretation at 30 or 24 fps. If you shoot 120 fps for a slow-motion moment, keep the rest of the edit at the project base rate so the speed change reads as intentional.
Whatever you choose, keep the timeline consistent. Mixing 24, 30, and 60 fps clips without conforming them creates judder that no amount of grading will hide.
Compression: the quiet reason your footage looks soft
What happens after you tap publish
Your upload is decoded, scaled, and re-encoded with settings tuned for storage cost and streaming. That second encode is lossy and it is not gentle with detail. Fine grain becomes blotchy, subtle gradients band, and thin high-contrast lines — text, wire fences, striped shirts — develop shimmer. The platform also generates multiple variants, so a clip may be served at a lower bitrate on a weak connection. You cannot control the final encode, but you can hand it an easier input.
Export settings that survive re-encoding
Aim for these habits:
- Codec: H.264 in an MP4 or MOV container for maximum compatibility. High profile, level 4.2 or higher.
- Bitrate: 10–16 Mbps for 1080p, 35–60 Mbps for 4K. Constant quality or a high variable bitrate beats a low fixed one.
- Keyframes: every 1–2 seconds. Dense keyframes reduce artefacts around fast cuts.
- Colour: deliver Rec. 709 with a legal range. Uploading flat log footage without conversion makes everything look grey and washed out.
- Sharpening: none, or very light. Sharpening halos amplify into ringing after a second encode.
- Noise reduction: apply sparingly. Over-denoised footage turns into plastic, and residual noise compresses badly.
If you have the option, avoid exporting far above the publish target just in case. Deliver a clean, correctly sized master and let the platform do less work.
Audio that keeps people watching
Sample rate, bitrate, and loudness
Audio quality drives perceived quality more than most creators admit. A sharp image with hollow, clipping sound feels cheap; a modest image with clear, well-levelled audio feels professional.
Deliver 48 kHz stereo, AAC at 256–320 kbps. Target around −14 LUFS integrated loudness with a true peak no higher than −1 dBTP, which keeps your mix competitive without triggering the limiter that protects listeners. Keep dialogue centred and consistent between takes; level jumps are more distracting than a slightly dull tone.
Designing for muted autoplay
Most feed views start muted. That means your first two seconds must work silently: a strong visual hook, on-screen text that states the premise, and motion in the frame. Treat captions as a design element rather than a compliance chore. Burn-in subtitles in a legible weight, placed inside the safe zone, outperform auto-captions that drift and mistime.
Music deserves care too. Use tracks from licensed libraries and let the edit follow the beat, because cuts that land on the downbeat feel deliberate. Duck the music 6–10 dB under voice, and check the mix on a phone speaker, since that is where most of your audience will hear it.
Safe zones, text placement, and cover frames
Interface overlays vary by placement and change over time, so think in bands rather than exact pixels. For vertical video, keep the bottom 20–25 percent clear of essential text, keep the top 10–12 percent clear, and keep roughly 6 percent of margin on each side. Anything outside that inner rectangle may be covered by a handle, caption, sticker, or button.
Cover frames matter more than most people assume. In a feed grid, the first frame is a thumbnail; on a profile, it is part of a visual system. Choose a frame with a clear subject, readable contrast, and no mid-blink expression. If your edit does not naturally start on a strong frame, set a cover deliberately.
One more framing note: faces and hands carry attention. Placing a talking head slightly off-centre, with the eyeline in the upper third, gives you room for text without covering the subject.
A repeatable AI-assisted production workflow
Step 1 — Brief and script
Write a one-sentence premise and a beat sheet before touching any tool. Specify placement, aspect ratio, duration, and the single idea the viewer should remember. A generative model can produce beautiful footage with no narrative spine; the script is what keeps it coherent.
Step 2 — Generate, gather, and license footage
This is where modern generative video models earn their place. Text-to-video and image-to-video tools can create establishing shots, abstract backgrounds, stylised transitions, and product contexts that would be expensive to shoot. Image models help you lock a visual style, which then carries into motion. Upscalers and frame interpolation tools rescue generated clips that are slightly soft or slightly short.
Habits that save time:
- Generate short clips — three to five seconds each — and cut them into a sequence rather than asking for one long shot. Coherence and control both improve.
- Fix the look first. Choose lens, lighting, palette, and grain direction, then repeat those descriptors in every prompt.
- Generate more than you need. Expect a hit rate and select ruthlessly.
- Keep provenance tidy. Note which clips are generated and which are filmed, and follow platform disclosure expectations for synthetic media.
If you produce many variants, batch the work and keep a queue rather than babysitting single renders. A structured pipeline — prompt set in, rendered clips out, organised by shot — keeps a week of production from collapsing into scattered files.
Step 3 — Edit, stabilise, and grade
Assemble in an editor such as DaVinci Resolve, Premiere Pro, Final Cut, or CapCut. Conform all clips to one frame rate, stabilise where needed, and grade to a neutral Rec. 709 finish. Generated footage often has slightly different grain and contrast from shot to shot, so a unifying grade, a light grain pass, and consistent colour temperature do more for believability than any single clip.
Sound design belongs at this stage: room tone, whooshes, footsteps, and music. Silence between cuts is a common giveaway of synthetic footage.
Step 4 — Export presets per placement
Build presets so you never re-derive settings under deadline:
- Vertical preset: 1080 × 1920, 30 fps, H.264, 12 Mbps, AAC 320 kbps, 48 kHz.
- Feed preset: 1080 × 1350, 30 fps, same codec and audio settings.
- Square preset: 1080 × 1080, identical otherwise.
- Archive master: 2160 × 3840 ProRes or high-bitrate H.264 for future re-edits.
Name presets by placement rather than by date, and keep them in a shared folder if you work with others.
Step 5 — Quality check and publish
Watch the exported file on an actual phone before uploading. Check the first two seconds muted, the last frame, the loudness, and caption placement. Then upload, and check the published version again, because platforms occasionally shift crops and overlays.
Pre-publish checklist
- Aspect ratio matches the placement, with no unintended letterboxing.
- Essential text sits inside the inner safe rectangle; nothing important in the bottom fifth.
- Master resolution is clean, with no accidental upscaling.
- Frame rate is consistent across all clips.
- Bitrate is high enough to survive re-encoding, with no heavy sharpening.
- Audio is 48 kHz stereo, around −14 LUFS, with true peak under −1 dBTP.
- Captions are burned in or verified and readable on a small screen.
- Cover frame is deliberate and legible.
- Music and footage licensing is documented.
- Synthetic-media disclosure handled where required.
Common mistakes and how to fix them
A few failures show up again and again.
Letterboxed horizontal video in the feed. Fix by reframing to 4:5 or 9:16, or by designing a deliberate background rather than accepting black bars.
Text hidden behind interface elements. Fix by moving text into the middle band and testing on a phone.
Soft footage after upload. Usually a bitrate or sharpening issue, or a master that was upscaled. Export a clean 1080p file at 12 Mbps and avoid over-processing.
Juddering motion. Almost always a frame rate mismatch during editing. Conform everything to the project rate before cutting.
Audio that clips or jumps. Normalise dialogue per clip, then apply a single loudness target across the timeline.
Generated clips that feel disconnected. Fix with a shared grade, consistent lens language, and a sound bed that connects shots.
FAQ
What resolution should I upload? Publish at 1080 pixels on the short edge — 1080 × 1920 for vertical, 1080 × 1350 for feed. Master at 2160 × 3840 when your workflow allows, so you have room to stabilise and reframe.
Is 60 fps better than 30 fps? Not inherently. Thirty frames per second is the safest default for most content. Use 60 for fast motion or slow-motion flexibility, and 24 for a filmic look.
Why does my video look worse after upload? Because the platform re-encodes it. Hand it a clean, correctly sized master with a generous bitrate, minimal sharpening, and controlled noise, and the second encode has less to destroy.
Should captions be burned in? For feed autoplay, burned-in captions are more reliable and can be styled to match your brand. Auto-captions are convenient but drift, mistime, and occasionally mishear.
How long should a Reel be? As long as the idea needs and no longer. Fifteen to thirty seconds suits most single-idea pieces; longer works when there is a genuine narrative arc or a tutorial structure.
Can AI-generated footage hold up in a real campaign? Yes, particularly for backgrounds, establishing shots, conceptual visuals, and effects. The failure mode is a string of unrelated pretty shots, which is a scripting and editing problem rather than a generation problem.
What about loudness differences between platforms? Normalise to about −14 LUFS integrated for social delivery and keep true peaks below −1 dBTP. That leaves headroom for platform loudness normalisation without obvious pumping.

