Start Free Now
Limited Time Offer: Get 50% OFF Starter & Basic Yearly Plans 🎉

AI Video Upscaling and Compression: An FFmpeg Workflow Guide

Sep 27, 2026

Start With the Delivery Target, Not the Tool

Most disappointing exports are not the fault of a weak encoder or an old camera. They come from deciding on resolution, bitrate, and enhancement strength before anyone asked where the file would actually be watched. A creator who commits to 4K and aggressive sharpening before checking the destination ends up with a master that is slower to render, heavier to store, and softer once a platform re-encodes it.

Working backwards fixes this. Answer three questions first.

Where will this be viewed? A phone, a laptop, a living-room television, and a projector fail in different ways. Phone screens hide blocking but expose banding across gradients. Large televisions expose everything, especially noise and edge halos.

What will the platform do to the file? Nearly every distribution platform re-encodes uploads using its own ladder, its own color handling, and its own quality targets. A file that looks flawless on your timeline can lose its fine detail during that second encode. Understanding this changes how much headroom you should leave.

How long must the file survive? A social clip lives for days. An archive master must survive for decades, which changes how aggressively you can compress it.

Once those answers are fixed, technical choices become comparisons instead of guesses. Resolution becomes a requirement rather than a badge. Bitrate becomes a budget to allocate. Enhancement becomes a question of how much the source can support.

This guide covers a repeatable pipeline: inspect the source, repair damage, upscale with a model matched to the content, encode perceptually, re-frame for multiple aspect ratios, and verify against the original. It is written for editors, post-production generalists, and technical creators who want one workflow that handles both archive rescue and fresh camera material.

Three Concepts People Constantly Merge

Almost every argument about video quality is really three arguments tangled together.

Resolution is not detail

Scaling 720p footage to 4K with a good resampler produces a larger frame holding exactly the same information. It looks smoother, not sharper. Super-resolution models change that equation by predicting high-frequency structure that plausibly belongs in the scene: individual hair strands, the weave of fabric, the crisp edge of lettering on a sign. That prediction is educated guesswork, which is why the same model can be remarkable on one clip and unsettling on the next. It can rebuild a face convincingly while inventing a pattern in a brick wall that was never there. Rule of thumb: upscale when the destination genuinely needs the pixels and the footage is clean enough to give the model real signal. A careful 1.5x or 2x pass on well-lit, low-noise material beats an aggressive 4x pass on anything.

Bitrate is a budget, not a score

A 20 Mbps file can look worse than a 6 Mbps file when the bits are spent in the wrong places. Allocation matters more than volume: how many bits a fast pan receives, how many a static interview shot receives, and whether the encoder understands what a viewer will notice. A talking head against a plain wall needs very little data. A handheld walk through foliage needs a great deal.

Perception is the referee

Every codec and every enhancement model is built on a model of human vision: contrast sensitivity, masking inside busy texture, tolerance for motion blur. Perceptual coding spends data where the eye is looking and saves it where the eye is not. That is why a well-tuned encode at half the size can be indistinguishable from a bloated one, and why a badly tuned encode at the same size can look obviously broken.

Decision rule: if you can only improve one variable, improve the encode. Encoding discipline affects every second of the file, while a modest upscale mostly affects how the picture holds up at full magnification.

Inspecting the Source Before You Touch a Pixel

Skipping measurement is how a project produces a supposedly better file that is quietly softer than the original. Ten minutes of inspection saves hours of re-rendering.

Read the streams first

Probe the file without decoding a single frame. ffprobe reports codec, profile, pixel format, frame rate, bit depth, color primaries, transfer characteristics, and matrix coefficients. A surprising number of washed-out and oversaturated complaints trace back to a mismatch between BT.709 and BT.2020 tagging rather than to compression. If the metadata is wrong, fixing it costs nothing. If you re-encode without fixing it, the error becomes permanent.

ffprobe -v error -show_streams -show_format input.mp4

Metrics that actually guide decisions

  • PSNR is cheap and universally reported, but it correlates poorly with perceived quality. Use it as a sanity check on whether something went catastrophically wrong, not as a verdict.
  • SSIM models structural similarity and sits in the middle: more perceptually meaningful than PSNR, less accurate than VMAF.
  • VMAF is trained on human ratings and is the best general-purpose default for comparing an encode against a reference.
ffmpeg -i encoded.mp4 -i reference.mp4 \
  -lavfi libvmaf='log_fmt=json:log_path=vmaf.json' -f null -

A habit that pays off: generate three encodes at different quality targets, score each one, then watch all three on the smallest screen you care about. Metrics catch regressions, while eyes catch wrongness. When the two disagree, trust your eyes and investigate the metric later.

Restoration Before Enhancement

Enhancement amplifies whatever is already in the frame, including damage. Restoration has to come first, and it has to be gentle enough to preserve real texture.

Reading the damage signature

Blocking appears as square boundaries in flat areas and shadows. Mosquito noise appears as shimmering specks around sharp edges and text. Banding appears as stair steps across gradients. Each problem has a different remedy, and applying the wrong one wastes time while erasing detail you cannot get back.

ffmpeg -i source.mkv -vf 'deblock=filter=strong,hqdn3d=2:1:3:3' \
  -c:v libx264 -crf 16 -preset slow -pix_fmt yuv420p10le repaired.mkv

For heavier restoration, a non-local means denoiser is far more thorough and dramatically slower. Reserve it for short clips that matter: a title sequence, a client testimonial, or a hero shot inside a montage. For everything else, a light spatial-temporal denoise is the better trade.

Always denoise before upscaling

This is the single highest-leverage habit in the entire workflow. Models treat noise as signal. Upscale a grainy clip and the grain becomes sharper, larger, and far more expensive to encode. A clip that needed 6 Mbps before enhancement can demand 20 Mbps afterwards with no visible improvement to show for it.

When restoration is not worth it

Not every source deserves rescue. A five-second clip buried in a fast montage does not justify an overnight render. A shot whose real problem is a focus miss cannot be fixed by any model — sharpening a soft frame produces crisp blur, which reads worse than the original. Be willing to declare a shot unrecoverable, crop it, cut it, or use it smaller in frame.

Matching an Upscaling Model to the Material

Content categories and model families

  • General photographic models handle live action competently and are the safest default for interviews, documentary footage, and product shots.
  • Animation-tuned models understand flat color regions, clean outlines, and limited palettes. On cartoons they routinely beat general models, which tend to add unwanted micro-texture to flat areas.
  • Restoration-focused models trained on compression artifacts can rescue a badly encoded archive clip that a pure upscaler would turn into a soup of invented texture.
  • Face-aware models improve skin and eyes but can over-smooth. Use them at partial strength when the rest of the frame matters.

Selection criteria, in priority order: source quality (clean, high-bitrate footage tolerates strong upscaling, while blocky low-bitrate footage needs artifact removal first); content type (faces, text, and architecture benefit most, while grass, water, and smoke are where models most often hallucinate); temporal behavior (per-frame processing flickers on fine detail); and degradation realism (training on real-world noise usually beats training only on synthetic downscaling).

The flicker problem

Processing each frame in isolation produces shimmer that is far more distracting than mild softness. Three practical mitigations: use models or pipelines that accept several neighboring frames as input; process in overlapping windows and blend the boundaries so discontinuities disappear; and post-stabilize with a light temporal denoise that smooths frame-to-frame variation without killing genuine motion.

ffmpeg -i repaired.mkv \
  -vf 'dnn_processing=dnn_backend=openvino:model=upscale_model.xml' \
  -c:v libx264 -crf 16 -preset slow -pix_fmt yuv420p10le enhanced.mkv

Backend availability changes between builds and platforms, so verify that your installation exposes the filter before designing a pipeline around it. When it works, it removes the need to move files between applications, which matters enormously for batch work.

Choosing a scale factor

A 1.5x pass that preserves texture usually beats a 4x pass that invents it. If delivery is 1080p and the source is 720p, upscaling to 1440p and letting a high-quality resampler finish is often cleaner than jumping straight to 4K. Test the hardest shot — fast motion, fine texture, a face — before processing the entire timeline.

Encoding With Perception in Mind

Codec comparison

Codec Efficiency Speed Compatibility Typical use
H.264 Baseline Very fast Universal Social delivery, review copies
HEVC / H.265 About 30% better Moderate Good, licensing-sensitive 4K delivery, archives
AV1 About 50% better Slower, improving Growing hardware support Web delivery, long-term masters
VP9 Between HEVC and AV1 Moderate Strong in browsers Web video where AV1 support is uncertain

The decision is rarely which codec is best and almost always which codec suits this destination. A universal H.264 file plus a high-efficiency AV1 master covers nearly every scenario you will encounter.

Quality targets versus hard ceilings

Constant quality is the sane default because it targets a level and lets the bitrate land wherever it needs to. Two-pass encoding makes sense when you must hit a hard size or bandwidth ceiling, because it lets the encoder plan allocation across the whole timeline instead of guessing.

ffmpeg -i enhanced.mkv -c:v libx265 -crf 22 -preset slow \
  -pix_fmt yuv420p10le -tag:v hvc1 -c:a aac -b:a 192k delivery.mp4

ffmpeg -i enhanced.mkv -c:v libsvtav1 -crf 30 -preset 6 -g 240 \
  -c:a libopus -b:a 128k delivery_av1.mkv

Two details matter more than people expect. First, a longer keyframe interval reduces overhead but costs quality at seek points; something in the five-to-ten-second range is a solid compromise for most web delivery. Second, 10-bit encoding improves compression efficiency even for 8-bit sources by reducing banding in gradients. It is one of the few genuinely free wins available.

Tune for the content, not the average

Grain is expensive. Dark scenes are expensive too, because shadow detail hides in low-amplitude values where quantizers bite hardest. A quality-targeted encode handles both gracefully. A fixed-bitrate encode will produce a blocky night scene and then waste bits on a bright, simple one.

A Step-by-Step Pipeline That Holds Up

This sequence works for most archive material and camera originals. The order is not arbitrary — changing it almost always costs quality, time, or both.

Conform and normalize

Decide frame rate, color space, resolution, and audio layout before any enhancement. Normalizing first prevents wasted work and stops small mismatches from compounding.

ffmpeg -i source.mov -c:v libx264 -crf 16 -preset slow -pix_fmt yuv420p10le \
  -c:a pcm_s16le conformed.mkv

Repair compression damage

Apply deblocking and deringing as gently as the artifact requires. Over-filtering flattens texture into plastic, and that damage cannot be undone later.

Denoise with restraint

Start weak and increase slowly, checking a static shot and a motion shot after each change. If you cannot see the noise anymore on a large display, you have probably gone too far.

Upscale, then re-sharpen carefully

Add noticeably less unsharp masking than feels right while you are editing, because enhancement reads stronger on a large screen at full brightness.

ffmpeg -i repaired.mkv \
  -vf 'scale=3840:2160:flags=lanczos,unsharp=5:5:0.6:3:3:0.2' \
  -c:v libx264 -crf 16 -preset slow enhanced.mkv

Consider frame interpolation last

Doubling the frame rate can help sports, screen recordings, and animation. It can also create soap-opera motion and warping around fast-moving limbs. Test it on the hardest shot in the timeline before committing the whole edit.

ffmpeg -i enhanced.mkv \
  -vf 'minterpolate=fps=60:mi_mode=mci:mc_mode=aobmc:vsbmc=1' \
  -c:v libx264 -crf 16 -preset slow motion.mkv

Encode for delivery

Produce at least two outputs: a high-efficiency master and a broadly compatible delivery file. Normalize loudness on the audio path so nobody reaches for the volume control.

ffmpeg -i enhanced.mkv -c:v libx264 -crf 20 -preset slow \
  -af loudnorm=I=-14:TP=-1.5:LRA=11 -c:a aac -b:a 192k final_web.mp4

Verify against the original

Check the three places where compression fails first: the opening ten seconds, the darkest scene, and the shot with the most motion. Watch the output and the source side by side at full size before you ship anything.

Re-framing and Multi-Format Delivery

Distribution now requires several aspect ratios from one shoot. Vertical, square, and horizontal versions are standard, and blind center-cropping is the fastest way to lose the subject of a shot.

  1. Identify the subject in each shot: a face, a hand, a moving object. Shots with no clear subject can be cropped for composition instead.
  2. Define a safe corridor for the crop so small tracking errors never clip a head or a hand.
  3. Smooth the crop path. Automatic tracking that snaps between positions looks worse than a manual keyframe with a slow ease.
  4. Rebuild headroom. Vertical crops almost always need the subject repositioned lower or off-center to leave space above.
  5. Re-check the encode. A tighter crop means fewer pixels carrying the same detail, which often requires a stricter quality target.
ffmpeg -i horizontal.mkv -vf 'crop=ih*9/16:ih,scale=1080:1920:flags=lanczos' \
  -c:v libx264 -crf 19 -preset slow -c:a copy vertical.mp4

Plan for overlays too. Platform interfaces cover the top and bottom of vertical video, so keep critical framing away from the extremes.

Hardware, Batching, and Scheduling Discipline

GPU encoders are dramatically faster and slightly less efficient at the same quality target. For review copies, social delivery, and long batch jobs, that trade is usually correct. For final masters where every megabyte counts, CPU encoding at a slow preset still wins on quality per bit.

Two rules keep batch work sane. Never run an expensive enhancement pass across an entire timeline before the edit is locked: work with proxies, finalize the cut, and process only what survives. And write scripts that log every parameter that produced a file, keep source files read-only, and always process one short test clip before launching a queue.

Budget enhancement by shot, not by timeline. Hero shots and close-ups justify heavy processing. Background montage footage rarely does. Consistency matters more than peak quality on any individual frame: a clip that is sharp for ten seconds and soft for two looks broken, while a uniformly good clip looks intentional.

Common Mistakes and How to Avoid Them

  1. Upscaling before denoising. Noise becomes permanent texture and inflates the file size at the same time.
  2. Sharpening twice. Once in the enhancement stage and again in the encoder tune, edges turn into visible halos.
  3. Ignoring color metadata. A tagged mismatch makes a technically perfect encode look dull or oversaturated.
  4. Judging on a small preview. Compression artifacts and flicker hide in thumbnails and reveal themselves on a television.
  5. Re-encoding repeatedly. Every generation loses a little. Keep lossless intermediates between stages.
  6. Chasing maximum resolution. Producing 4K for a platform that serves 1080p wastes time and can look worse after the platform re-encodes it.
  7. Over-filtering faces. Skin loses pore-level detail and reads as plastic, which is far more distracting than a little noise.

A quick diagnostic habit: when something looks wrong, change exactly one variable and re-render a thirty-second section. Adjusting three settings at once tells you that something improved, but not what.

FAQ

Why does my upscaled video look waxy?

Typical causes are a photographic model applied to animation, an overly aggressive denoise before upscaling, or sharpening layered on top of already reconstructed edges. Reduce denoise strength first, then lower the scale factor.

Why does the file look fine on my laptop but banded on a television?

Gradients are the most fragile part of any encode, and larger screens with different gamma curves expose them. Encode in 10-bit, raise quality slightly for dark scenes, and consider a mild dithering pass when banding is severe.

Should I always upscale to the highest available resolution?

No. Upscale only to what delivery actually needs. Extra pixels cost render time and storage, and they can vanish anyway in a platform re-encode.

How do I compare two encodes fairly?

Use the same source, the same duration, and the same playback conditions. Compute a perceptual metric, then watch both at full size on a large display. If you cannot tell them apart at a normal viewing distance, take the smaller file.

Can I run this without a GPU?

Yes, but scale your expectations. CPU-only pipelines are viable for short clips and moderate resolutions. For long-form material, a GPU turns an overnight job into an hour.

When should I stop enhancing?

When the output is stable, clean, and faithful to the original performance. A clip that looks slightly soft but natural almost always reads better than one pushed until it looks artificial.

Does frame interpolation belong in every project?

No. It helps sports, screen capture, and animation. On narrative footage it frequently produces motion that reads as artificial, especially around hands and fast turns.

What is the best order for the whole pipeline?

Repair, denoise, upscale, interpolate, encode, verify. Changing the order almost always costs quality or time, and often both. The goal is not maximum sharpness. It is a delivery file that respects the source, survives a second encode, and looks correct on whatever screen the audience happens to be using.

Alexander

Alexander