Start Free Now
Limited Time Offer: Get 50% OFF Starter & Basic Yearly Plans 🎉

AI Video Restoration: Make Old Footage Sharp Again

Oct 6, 2026

Why Old Footage Falls Apart — And What AI Can Actually Fix

Every archive tells the same story. A tape sits in a drawer for twenty years, gets digitized once at whatever settings the capture card defaulted to, and then spends another decade as a 480p file with blocky compression artifacts, wobbling scan lines, and audio that drifts out of sync by the three-minute mark. When someone finally asks to use it, the footage looks nothing like what the people in it remember.

AI restoration changes that equation, but not in the way marketing pages suggest. The technology does not "unlock" hidden detail that was never recorded. What it does is far more interesting: it learns what plausible detail looks like at higher resolutions, then synthesizes it in a way that is consistent with the motion, lighting, and texture of the original frame. Done well, the result reads as a sharper version of the same footage. Done badly, it reads as a completely different film with a stranger's face pasted onto your grandmother.

Understanding that boundary is the whole game. Restoration breaks into four distinct problems, and each one needs a different kind of model:

  • Signal loss — compression blocking, tape noise, and analog grain that obscure real detail.
  • Resolution loss — a 720x480 frame has to become 3840x2160 without turning into mush.
  • Temporal loss — dropped frames, duplicated frames, and juddering motion.
  • Perceptual loss — the soft, slightly muddy look that makes viewers say "this feels old" even when technically nothing is broken.

A good pipeline addresses these in order. A bad pipeline throws everything at a single upscaler and hopes. The rest of this guide is the ordered version.

Assess the Source Before You Touch a Model

The single most common mistake in AI restoration is skipping assessment and going straight to a 4x upscale. Ten minutes of inspection saves hours of rendering.

Inventory format, codec, and cadence

Start by writing down hard facts about every clip:

  • Original medium: VHS, Hi8, MiniDV, 16mm film scan, DVD rip, or an early digital file.
  • Container and codec: MPEG-2, DV, H.264 at a low bitrate, or ProRes from a film scanner. Codec matters because blocky MPEG-2 artifacts need different treatment than smooth analog grain.
  • Resolution and pixel aspect ratio: many SD formats used non-square pixels. If you upscale before correcting the aspect ratio, faces stretch permanently.
  • Frame rate and scan type: 23.976, 25, 29.97 interlaced, or a variable frame rate from a phone screen recording.
  • Field order: upper or lower field first. Getting this wrong produces a visible comb pattern on every motion frame.
  • Audio: sample rate, channel layout, and any drift.

Put this in a spreadsheet. You will reference it constantly, and it doubles as delivery documentation.

Decide the target before you render

"Make it look better" is not a target. Choose deliverables up front:

Use case Typical target Notes
Family archive 1080p, original cadence Prioritize stability over sharpness
Documentary insert 1080p or 2160p, 23.976 or 25 Must cut against modern footage
Social clip 1080x1920 vertical Crop decisions happen before upscale
Broadcast delivery 1080i or 1080p per spec Loudness and safe-area rules apply
Museum / cultural archive Preservation master at highest feasible res Keep an untouched mezzanine

Always keep an untouched mezzanine file. Every generative step is destructive in the sense that it commits to an interpretation. If a reviewer later says the faces look wrong, you need a clean source to return to.

Deinterlace and correct aspect ratio first

Interlaced material must be deinterlaced before any AI model sees it. Generative upscalers interpret comb artifacts as texture and will happily sharpen them into permanent horizontal streaks. Use a proper motion-adaptive deinterlacer, not a simple blend. Likewise, correct pixel aspect ratio and any geometry distortion (lens warp, tape head skew) before upscaling.

The Restoration Pipeline, Stage by Stage

Treat restoration as a chain where each stage feeds a clean plate to the next. Order matters more than tool choice.

Stage 1: Stabilize and repair

Fix physical and mechanical problems first: gate weave, vertical jitter, dropped-frame stutter, and torn frames. A stabilizer that smooths handheld motion is different from one that removes mechanical jitter — use the latter, and set it conservatively. Over-stabilization creates a floating, rubbery frame that no later stage can fix.

For dropouts and scratches, patch them here. A scratch that survives into the upscale stage becomes a sharp, confident white line.

Stage 2: Denoise and degrain

This is where most quality is won or lost. Noise is high-frequency data, and every upscaler treats high-frequency data as detail. Feed it grain and you get grain at 4K, crisp and permanent.

A practical approach:

  1. Run a temporal denoiser first. Temporal models compare adjacent frames and remove noise that is not consistent across time, which preserves real detail far better than spatial-only denoising.
  2. Follow with a light spatial pass only where needed. On grainy film, keep some grain — a fully degrained frame looks plastic and forces the upscaler to invent texture.
  3. Never denoise twice at strong settings. Stacking produces waxy skin and smeared foliage.

For analog tape, add a chroma-noise pass. VHS chroma bleeding is a separate problem from luma noise and needs its own correction or color will smear across edges.

Stage 3: Upscale and synthesize detail

Now, and only now, apply the upscaler. Two families of models dominate:

  • Restoration-trained upscalers are trained on degraded-source pairs, so they are conservative and tend to preserve original character. Best for archival work.
  • Generative diffusion upscalers hallucinate plausible detail. They produce breathtaking results on landscapes and architecture, and uncanny results on faces and text.

A hybrid strategy usually wins: conservative model for the base upscale, then a light generative pass at low strength to restore micro-texture like fabric weave, hair strands, and skin pores.

Work in tiles or segments rather than the entire timeline at maximum strength. Render a 10-second representative clip, watch it, and only then commit to full-length processing.

Stage 4: Face and texture recovery

Dedicated face restoration models are powerful and dangerous. They can reconstruct plausible eyes and teeth in a blurry shot, but they also impose a generic identity. Use them per-shot, with the strength dialed back, and compare frames side by side against the source at 200% zoom. If a subject's face changes shape between shots, you have gone too far.

For group shots and crowd scenes, prefer a general upscaler with mild face enhancement over aggressive per-face reconstruction. Consistency beats sharpness.

Stage 5: Frame interpolation and cadence

If the source is 12 or 15 fps — common in early film and animation — frame interpolation can smooth motion. But interpolation invents frames, and it invents them badly around fast motion, occlusion, and cuts. Two rules:

  • Interpolate to a clean multiple (12 to 24, 15 to 30), never to an arbitrary number.
  • Mask out cuts and fast-motion segments, or handle them with optical-flow-aware settings.

When in doubt, preserve the original cadence and let the projector or player handle it. Authentic judder is often more pleasing than smooth artifacts.

Stage 6: Color, grain, and finishing

Restore color last. Generative models can shift color subtly, so grade after upscaling, not before. Rebuild black levels, correct any tape-induced color cast, and then decide on grain. A final light grain layer matched to the target format helps AI-restored footage sit naturally inside a modern edit rather than looking glassy.

This is also where you normalize audio, which is covered below.

Choosing the Right Model for Each Job

Model choice should follow the problem, not the hype cycle.

Match the model to the defect

  • Heavy compression blocking: use a restoration-trained model with strong deblocking, ideally one trained on streaming or broadcast degradation.
  • Analog tape noise: temporal denoise plus a conservative upscaler; generative models amplify tape artifacts into texture.
  • Film grain and gate weave: film-specific restoration models handle grain structure better than generic upscalers.
  • Animation and line art: restoration-trained models preserve clean edges; generative models tend to warp line weight.
  • Text overlays, signage, lower thirds: use the most conservative model available. Any generative pass will produce plausible-looking gibberish.

When a diffusion model helps — and when it hurts

Diffusion-based video upscalers shine when the scene has strong structural priors: bricks, foliage, crowds at a distance, water, architectural detail. They fail on the things viewers are most sensitive to: faces, hands, logos, printed text, and anything the audience knows by heart.

A useful heuristic: if a viewer will recognize the subject personally, be conservative. If the subject is anonymous texture, generative detail is free real estate.

Run a bake-off on three shots

Before committing, pick three representative shots — one close-up face, one wide landscape, one motion-heavy action beat. Run every candidate model on those three. Score them on stability (no flicker), identity fidelity (does the person still look like themselves), and texture plausibility. The winner is usually not the model that looks sharpest on a still frame.

Managing Compute, Storage, and Long Renders

Restoration is a compute problem as much as a creative one, and plans collapse when a 90-minute documentary meets a 4x upscale.

Estimate before you render. A rough rule: processing time scales with output pixels, model size, and temporal window. If a 10-second test takes four minutes, a 90-minute reel is not going to finish overnight on the same machine.

Chunk long timelines. Render in 5- to 10-minute segments. Chunking gives you restart points, lets you spot drift between segments, and makes it possible to use multiple machines. Just overlap by a few frames so you can match the seams.

Watch for temporal flicker between chunks. Generative models can drift in tone and detail across a long render. If segment 3 looks slightly cooler than segment 2, you will see a visible pop at the cut. Grade segments against each other, or use a fixed seed and reference frame where the tool allows it.

Store intermediates as high-bitrate, near-lossless files. Re-encoding between stages compounds compression artifacts and partially undoes your work. Disk space is cheap; a ruined master is not.

Keep a render log. Model, version, settings, strength values, and the seed for every segment. When a client asks for one shot to be re-done more softly six months later, the log is the difference between a ten-minute fix and a full re-render.

Audio Restoration: The Half Everyone Forgets

A perfectly restored picture with untouched audio still feels old. Listeners forgive soft footage far more readily than hiss, hum, and clipping.

A standard audio pass:

  1. Remove hum and rumble with a narrow notch filter at the mains frequency and a gentle high-pass around 60–80 Hz.
  2. Denoise with a learned model, not a broadband gate. Modern speech-enhancement models separate voice from steady noise far better than expanders, and they do it without the pumping artifacts that make dialogue sound underwater.
  3. Repair clicks, crackle, and dropouts with a declicker, then check for any transient damage the model introduced.
  4. Fix sync. Analog captures often drift. Measure drift at the head and tail, then correct with a continuous stretch rather than periodic nudges.
  5. Match loudness to the delivery spec, and check mono compatibility if the material may play on a phone speaker.

When audio is unsalvageable, do not be afraid to replace it. Room tone, ambience, and foley from a library will carry a restoration further than aggressive processing of a ruined track.

Quality Control: Checklists and Common Failure Modes

Watch restored footage at normal speed on a real screen, not just scrubbing through frames. Many artifacts only appear in motion.

The QC checklist:

  • Faces: identity preserved across every shot? No shape-shifting between cuts?
  • Text: signage, titles, and overlays still legible and correct?
  • Edges: no halos, no ringing, no sharpening outlines around high-contrast subjects?
  • Motion: no warping on fast pans, no melting around occlusions, no ghost frames?
  • Texture: skin, fabric, and foliage plausible at 100% zoom?
  • Consistency: no brightness, color, or sharpness pops between segments?
  • Audio: sync holds from head to tail, no artifacts on sibilants?
  • Cadence: no duplicated or blended frames on motion?

Common failure modes and their causes:

  • Waxy skin — over-denoised before upscale, or generative face model at high strength.
  • Crawling texture — temporal instability, usually from frame-by-frame processing without temporal awareness.
  • Ghosting trails — interpolation applied across cuts or heavy occlusion.
  • Flickering grain — grain added after upscale without temporal coherence.
  • Smeared chroma — chroma denoising applied too aggressively, or skipped entirely on tape sources.
  • Detail that vanishes on close inspection — the upscaler invented structure that dissolves at 200% zoom. Acceptable for social, not for archive.

Three Practical Workflows

Family archive: one tape, one evening

Deinterlace, correct aspect ratio, run temporal denoise at moderate strength, upscale 2x with a conservative restoration model, apply a mild face pass at low strength, grade, add light grain, normalize audio. Deliver 1080p H.264 plus a lossless master. Runtime on a modern GPU: roughly two to four times real time.

Documentary insert: matching modern footage

Work at 1080p or higher with the original cadence preserved. Use the conservative upscaler for the base and reserve generative detail for wide shots only. Grade to match the surrounding modern footage — restoration that looks cleaner than the modern material cuts badly. Add matched grain and check the insert on a calibrated monitor in context.

Vertical social clip: acknowledge the crop first

Decide the crop before upscaling, because the upscaler should work on final framing. Upscale the cropped frame, apply a strong face pass since the subject will be large in frame, stabilize aggressively because handheld footage reads worse in vertical, then burn in captions after restoration — never restore burned-in text.

Frequently Asked Questions

Does AI restoration add detail that was never there?
It synthesizes plausible detail, not recovered detail. The distinction matters for archival ethics: mark generative restoration clearly, and always retain an untouched master.

Can I upscale directly from a phone screen recording?
Yes, but variable frame rate is a problem. Convert to constant frame rate first, or temporal models will produce stutter and ghosting.

Should I denoise before or after upscaling?
Before, almost always. Upscalers amplify noise into permanent texture. A light repair pass after upscaling is fine, but the heavy lifting belongs upstream.

Why does restored footage look artificial even when it is sharp?
Usually because grain was removed entirely, or because generative detail was applied uniformly. Real footage has noise. Adding matched grain and varying model strength per shot restores the natural feel.

How do I handle 4:3 material for a 16:9 delivery?
Choose deliberately: pillarbox with a tasteful background, crop and accept the loss, or use a blurred-fill composite. Never let a model "extend" the frame unless you are prepared to supervise every shot.

Is restoration worth it for footage under one minute?
Often yes, because a single strong clip can anchor a whole project. But run the three-shot bake-off first so you know which model suits that particular source.

Wrapping Up

Good AI video restoration is not a button. It is a sequence of decisions: understand the defect, fix what is mechanically broken, remove noise without killing texture, upscale conservatively, synthesize detail only where it is safe, and finish with color, grain, and audio that respect the original.

The teams that get the best results are not the ones with the biggest models. They are the ones who test on short segments, keep a clean master, log every setting, and know exactly when to stop. Sharpness is easy to add and almost impossible to remove. Restraint is the real skill — and the difference between footage that feels revived and footage that feels like it was replaced.

Alexander

Alexander