Why AI Video Detection Became an Everyday Skill
A few years ago, spotting synthetic footage was a niche hobby for machine-learning researchers. Today it is part of ordinary work for editors, publishers, recruiters, teachers, fraud teams, and anyone who forwards a link. Text-to-video systems such as Sora, Veo, Kling, Runway, Pika, Luma Dream Machine, Hailuo, and Wan can turn one sentence into coherent motion, believable lighting, and lip-synced speech. The question is no longer whether synthetic video exists. The question is whether the clip in front of you came from a camera, a person, or a model.
The practical stakes are easy to underestimate. A brand that publishes a synthetic product demo without a label can lose customer trust within hours. A newsroom that airs a fabricated street scene damages its own reputation for months. A hiring team that treats a cloned video interview as genuine wastes time and opens a door to identity fraud. Meanwhile, a legitimate animator using generative tools can be wrongly accused of deception, which is its own serious harm.
So detection is not only about catching fakes. It is about making fair, defensible judgments from evidence you can actually gather. This guide walks through provenance data, visual forensics, motion and physics, audio signals, context, a repeatable verification workflow, the limits of automated detectors, common mistakes, and the questions people ask most often.
One mindset helps from the start: treat detection like building a case rather than flipping a coin. No single clue proves anything on its own. What matters is a pattern of independent signals pointing the same direction, plus a written record of how you reached your conclusion. If you cannot explain your reasoning in three plain sentences, you probably do not have enough evidence yet.
It also helps to separate two very different questions. The first is technical: was this footage synthesized or captured? The second is ethical: was the audience told? A clip can be fully synthetic and completely honest, or shot on a real camera and deeply misleading. Keeping those questions apart prevents most bad calls.
Start With Provenance: Metadata, Platform Labels, and Watermarks
Every media file carries a small history of itself. That history lives in two places: embedded metadata inside the file, and labels attached by the platform that hosts it. Provenance is the least glamorous layer of detection and the most useful, because it gives you facts rather than impressions.
On a desktop computer, tools such as ExifTool and ffprobe can read the container and stream data of any file you have downloaded. A genuine camera clip often includes the device make and model, lens information, exposure settings, ISO, a creation timestamp, and sometimes GPS coordinates. Synthetic output frequently contains none of that, or contains a generic encoder string instead. However, absence of camera data is weak evidence on its own, because messaging apps and social platforms routinely strip metadata during upload. Presence of consistent camera data is stronger evidence in the other direction, though it can also be forged by anyone patient enough.
A few metadata habits are worth building. Check the declared resolution and frame rate. Many generators emit unusual dimensions such as square or vertical ratios that do not match common camera sensors, or lock frame rates at clean values without the tiny drift real cameras show. Check the creation timestamp against the story being told; a clip presented as a live event but stamped weeks earlier deserves a second look. Check the file name and the encoder line, which sometimes reveal the pipeline that produced the file.
Provenance also includes signed content manifests. The C2PA standard, backed by the Content Authenticity Initiative, lets cameras, editing apps, and publishing platforms attach a tamper-evident record of how a file was made and modified. When a manifest is present and valid, it is among the strongest signals available. When it is missing, you have learned only that the chain is unverified, which is extremely common on social platforms.
Invisible watermarking is the third provenance layer. Some generators embed a pattern in the pixels themselves rather than in the file metadata, so the mark can survive re-encoding and screenshotting to a degree. Watermark detection is imperfect and not universally deployed, but when a detector confirms a mark, it is meaningful. Platform disclosure labels sit on top of all this: several major video platforms now label content that was altered or generated, either automatically or through creator disclosure. Screenshot the label when you see it, because labels sometimes disappear after re-uploads.
Visual Forensics: Small Details That Give Generators Away
Visual inspection is where most people start, and where most people go wrong. The trick is knowing which details still carry signal and which ones have stopped being reliable.
Hands and fingers were once the classic giveaway, and they still fail sometimes, especially when hands interact with objects. But current models render hands well in simple poses, so a clean hand proves nothing. Better tells live in the details that receive less training attention: teeth in profile, jewelry clasps, the arms of eyeglasses, the shape of ears as a head turns, buttons that melt into fabric, and skin that looks uniformly smooth like polished plastic.
Text is another productive area. Background signage, license plates, product labels, tattoos, and on-screen lower thirds often degrade into plausible-looking nonsense. Letters may be shaped correctly but form no real words, or a logo may have the right silhouette with scrambled interior detail. Real signage tends to be legible even when small, because it was photographed rather than imagined.
Lighting logic is more reliable than texture. Real scenes obey physics: two light sources produce two shadow directions, specular highlights sit where the light actually is, and shadows soften with distance in a predictable way. Synthetic footage often has shadows that point in inconsistent directions, highlights that float without a source, or reflections in windows and mirrors that do not match the subjects standing in front of them. Reflections are a particularly strong check because they require the model to maintain a second consistent view of the scene.
Depth and focus are worth a slow look too. A real camera has a limited depth of field, so a subject in focus should be surrounded by gradients of softness that follow distance. Synthetic footage sometimes shows a sharp subject against a background that is uniformly blurred regardless of distance, or edges that shimmer and crawl during motion. The same applies to sensor noise: authentic footage has grain that shifts slightly frame to frame, while generated footage can carry a fixed noise pattern that sits perfectly still while everything else moves.
Finally, look at faces during speech. Face replacement and full generation both struggle with the boundary between jaw, neck, and collarbone. Watch for a visible seam along the jawline, teeth that flicker between frames, a neck that changes width, hair that dissolves into a collar, or an ear that redraws itself when the head rotates. These artifacts are subtle at full speed and obvious at quarter speed.
For a concrete example, imagine a viral clip of a street performer at night. Pause on the puddle below the performer. In a real recording, the reflection shows the same motion with slight ripple distortion. In many synthetic clips, the reflection shows a different pose, a missing object, or nothing at all. That single frame can settle a case that a hundred others leave open.
Motion, Physics, and Time: The Strongest Human Signals
If you only have time for one layer of manual review, spend it on motion. Generators are trained to produce convincing stills and short bursts, and coherence tends to decay as a clip continues. Most synthetic footage therefore works in windows of a few seconds, with errors clustering at transitions.
Watch for morphing in the middle of a shot. An object may slowly change shape, a sleeve may lengthen, a chair may lose a leg, or a background door may move between cuts. Object permanence is another failure mode: items that leave the frame and return may come back slightly different, or a person may pass in front of an object that never reappears.
Physics checks give you a mental test suite. Pouring liquid should accelerate and splash with weight. Cloth should fold along believable creases and stop moving when the body stops. Smoke should drift and dissipate. Hair should behave as separate strands with inertia rather than as a single rubbery sheet. Crowds should show independent motion rather than a synchronized shuffle. When a model has to track many interacting objects at once, errors multiply.
Camera behavior is a strong tell in both directions. Real handheld footage has micro-jitter, small exposure shifts, and imperfect framing. Synthetic footage often glides with impossible smoothness, or performs camera moves that no operator could physically achieve, such as an unbroken push through a wall and out the other side. Focus pulls are another giveaway: a real focus pull lands on a specific plane, while generated focus sometimes softens everything or snaps abruptly.
Pacing is a behavioral clue hiding in plain sight. Creators who want to hide defects cut often. If a clip is a rapid sequence of two-second shots with speed ramps, cross-dissolves, and quick zooms, ask whether the editing serves the story or the errors. Long, uninterrupted takes are harder to fake and therefore more informative.
Decision criteria matter here. A single morph in an otherwise continuous, physically consistent two-minute clip is more likely a compression glitch or a video effect. The same morph repeating at every transition across multiple clips from one account is a pattern. Patterns outweigh single events.
Audio Clues That Most Viewers Never Notice
Sound is often ignored, which makes it a productive place to look. Voice cloning and speech generation have improved dramatically, but they still leave traces in breath, timing, and room acoustics.
Start with lip-sync drift. Play the clip at half speed and watch a hard consonant such as p, b, or m. In genuine recordings, lips close at the exact moment the sound begins. In manipulated video, the closure often happens a few frames early or late, and the error changes across a sentence because the audio and video were generated separately.
Breath is the next check. Real speakers inhale, swallow, and pause in irregular ways. Cloned voices often produce long, perfectly smooth sentences with no breaths, or breaths inserted at mathematically tidy intervals. Prosody can be flat, with every sentence ending on the same falling tone, or emotionally mismatched, where the words are angry but the delivery stays neutral.
Acoustics tell their own story. A recording made in a tiled bathroom should sound bright and echoey, while a recording made in a carpeted office should sound dry. Synthetic audio sometimes applies reverb evenly to everything, including the speaker, or leaves the voice completely dry while the background carries room tone. Listen for a sudden change in background hiss when the shot cuts, or ambience that continues unchanged across cuts that should have moved the microphone.
Plosives, sibilance, and clipping are useful too. Aggressive p sounds in real speech cause brief low-frequency thumps and sometimes clipping. Generated speech may render them too cleanly or smear them across neighboring syllables. High-frequency s sounds may ring unnaturally, and loud passages may show compression artifacts that do not match the visual distance from the microphone.
Music deserves a quick mention. Fully generated music can loop in ways that are too regular, change instrumentation without reason, or fade out with no musical logic. A mismatched music bed that starts mid-phrase is a sign the audio was assembled rather than recorded.
Context Signals: Uploader Behavior, Framing, and Narrative Pressure
By the time a clip reaches you, it has a social history. That history is often more revealing than any pixel.
Look at the account. How old is it? What did it post last month? Does everything on the channel have the same visual texture, the same presenter, and the same thumbnail style? Accounts that publish dozens of unrelated viral clips per week, with descriptions full of generic phrasing, are aggregators rather than sources. Aggregators frequently repost synthetic content without knowing or caring where it came from.
Look at the framing. Clips designed to bypass scrutiny usually arrive with urgency: a shocking claim, no names, no location, no time, no follow-up interview, and a caption that demands you share it before the story is taken down. Real reporting tends to include specifics that can be checked, such as a street address, an agency name, or a journalist who will answer questions.
Reverse searching is a fast, high-value step. Extract three distinctive frames and search each one. You are looking for the earliest known appearance and for any version that clearly has a camera behind it. Be careful with results: a clip that appears on many large accounts is only popular, not verified. Popularity and authenticity are unrelated, and platforms often amplify the same item across unrelated feeds.
Cross-check the claim against the world it describes. Weather, daylight, architecture, license plates, street signs, uniforms, and accents can all be inconsistent with the stated location and season. A dramatic flood said to be in one city may show mountains that city does not have. These mismatches are cheap to check and hard to fake accidentally.
One more signal deserves attention: emotional payload. Content engineered to produce outrage, awe, or fear performs well, and high-performing content gets shared before it gets examined. When you notice your own reaction spiking, slow down. That reaction is a cue to verify, not evidence of anything.
A Repeatable Verification Workflow
Ad hoc checking produces inconsistent answers. A short workflow produces defensible ones and takes less time than people expect.
Preserve the original
Download the file rather than only bookmarking the page, and take a full-page screenshot that includes the uploader name, timestamp, caption, and any platform label. Save all of it in one folder. If the post is deleted later, your evidence survives, and you can compare what changed if it is re-uploaded.
Pull the metadata
Run ExifTool and ffprobe on the preserved file. Record the container, codec, resolution, frame rate, bitrate, encoder string, and any camera data. Note explicitly whether a content manifest is present. Absence of metadata is a neutral finding, not a verdict, so phrase it that way in your notes.
Watch it slowly, twice
First pass at normal speed for the overall story and emotional arc. Second pass at quarter speed with sound off, watching only the pixels. Then a third pass with sound on and the screen dimmed or off, so you hear the audio without being influenced by the image. This separation is what makes audio problems visible.
Extract and magnify the hardest frames
Pull stills from moments with the most complexity: hands interacting with objects, faces in profile, reflections, background crowds, text in signage. Zoom to full resolution and step frame by frame across a one-second range. Compare the same region across several frames; artifacts that shift or flicker are more telling than artifacts frozen in one frame.
Check the audio separately
Look for lip-sync offset on hard consonants, missing breaths, flat prosody, and inconsistent room tone across cuts. If you hear something odd, slow the audio to three-quarter speed; timing errors become obvious at slower rates.
Search for the earliest copy
Run reverse searches on three distinctive frames and on one unusual phrase from the caption. Note the earliest date you find, and note who published it. If every copy traces back to accounts that repost untraceable content, treat the clip as unverified regardless of how polished it looks.
Run a detector as one input, not a verdict
Choose one or two automated detectors, run them, and write down the score along with the exact file you submitted and the date. Then set the score aside and let your manual evidence lead. Detectors are useful for triage at scale and unreliable as standalone proof.
Write down your reasoning
Before you publish, forward, or act on a conclusion, write three sentences: what you observed, what you could not verify, and how confident you are. Use hedged language that matches your evidence. If your finding is a coin flip, say so, and describe what additional evidence would settle it.
Four decision criteria keep this workflow honest. First, require at least two independent signals before making a claim in public. Second, never rest a conclusion on style alone, because style legitimately varies. Third, prefer positive evidence, such as a valid content manifest, over negative evidence, such as a missing camera tag. Fourth, when you cannot decide, label the clip as unverified and explain what you checked; that is a legitimate result, not a failure.
What Detection Tools Can and Cannot Do
Automated detectors have become genuinely useful, and they are routinely oversold. They take a video, analyze frames for statistical patterns left by generation models, and return a probability. Names to expect in this space include Hive AI, Reality Defender, Deepware Scanner, Illuminarty, and AI or Not, along with platform-level provenance and watermark detectors. Different services often disagree on the same file, which is normal rather than scandalous: they were trained on different data, for different models, at different times.
Their strengths are speed and volume. A team reviewing hundreds of clips per day can use scores to sort items into likely-real, likely-synthetic, and unclear, then spend human attention where it matters. Their weaknesses are systematic. Detectors routinely score animation, 3D renders, video-game capture, time-lapse footage, slow motion, heavy color grading, beauty filters, and upscaled archival film as synthetic. They also degrade badly on compressed, cropped, or re-encoded files, which describes almost everything on social platforms.
Adversarial editing matters as well. Cropping a frame to remove a watermark, changing the frame rate, adding mild noise, or re-encoding through a different codec can push a synthetic clip toward a real verdict. Conversely, a real clip that has passed through several compression cycles can drift toward a synthetic verdict. This is why a score is only one line in your notes.
A practical rule: treat scores above a high threshold as a prompt to look harder rather than as a conclusion, and treat scores near the middle as no information at all. Record the tool, the version or date, and the exact file when you write the number down, because a score without that context is not reproducible.
Pair tools with human review in a fixed order. Human review first, detector second, human review again to interpret the score against what you actually saw. Any workflow that inverts that order ends up arguing with a number instead of examining a video.
Common Mistakes, False Accusations, and Edge Cases
The fastest way to improve is to know how verification usually goes wrong.
Relying on one signal is the most common error. Clean hands do not prove authenticity, one morph does not prove synthesis, and a missing camera tag does not prove anything. Build a pattern.
Treating a score as proof is the second. Detectors return probabilities, and probabilities are not findings. Quote the score as supporting evidence, never as the finding itself.
Forgetting that real footage can look synthetic is the third. Modern cameras apply heavy noise reduction. Editors stabilize, denoise, upscale, and color grade. Someone may have used a face filter, a portrait mode, or a digital zoom. Older footage that has been upscaled for a documentary can trip every detector you own.
Ignoring hybrid content is the fourth. Many clips are neither fully real nor fully generated: a real presenter with a synthetic background, a real street with an inserted vehicle, a real interview with a cloned voice patch. Hybrid edits are the most common deceptive format precisely because they are hardest to characterize.
Failing to check audio is the fifth. Audio review catches a large share of manipulations that visual review misses.
Skipping evidence preservation is the sixth. If the post disappears before you document it, your conclusion becomes unrepeatable.
Letting urgency override process is the seventh. An unverified clip shared with a clear caveat causes far less harm than a confident accusation that turns out to be wrong.
Confusing parody, satire, and stylized animation with deception is the eighth. Plenty of creators work in obviously artificial visual styles and label their work clearly. Enforcement should focus on deception, not on aesthetics.
The final mistake is failing to name the original source when you republish something. Even when a clip is genuine, republishing without attribution spreads confusion and buries the only person who could answer questions about it.
Frequently Asked Questions
Can I be completely certain a video is AI-generated?
Not with today's tools, and rarely with perfect certainty at all. You can reach high confidence by stacking independent evidence: valid provenance data, consistent physics failures across multiple shots, audio timing errors, and an upload history that shows a pattern. Say high confidence rather than proof, and describe the signals you relied on.
Do all AI videos have watermarks or metadata labels?
No. Some generators embed invisible marks, some attach signed provenance manifests, and many do neither, or the marks are removed by cropping and re-encoding. Social platforms also strip file metadata during upload. Missing labels mean unverified, not synthetic.
Why do two detectors give opposite answers?
They were trained on different model outputs, with different preprocessing and different thresholds. Scores shift with compression, resolution, and frame rate. When tools disagree, treat the result as unclear and lean on manual forensics.
Are hands still a reliable giveaway?
Only in complex interactions. Simple hand poses are usually rendered correctly now. Look at how hands grip, pour, button, or hand something to another person, and check the surrounding objects for warping at the same moment.
How can I check a video quickly on a phone?
Pause on a reflective surface and compare the reflection with the subject. Watch a face in profile during speech for jaw and ear seams. Listen for missing breaths. Scroll the uploader's recent posts for a pattern of untraceable viral clips. In two minutes, that combination catches most obvious cases.
Is it illegal to publish AI-generated video?
Rules vary by country and situation, and deception with intent to harm is treated very differently from clearly labeled artistic use. The durable principle is disclosure: label synthetic or altered content, avoid implying that real people said or did things they did not, and keep records of how the material was created.
What should I do if my own video is wrongly accused?
Respond with evidence rather than argument. Share the project file, the raw footage with timestamps, camera metadata, and a description of what software you used. Content manifests and behind-the-scenes clips end disputes faster than explanations do.
How long should verification take?
A quick triage pass takes a few minutes and answers most everyday questions. A publishable investigation of a high-stakes clip can take an hour or more once you include provenance review, frame analysis, audio inspection, source tracing, and written notes.
What is the fastest single red flag?
Motion that decays. If a clip looks convincing for three seconds and then morphs, warps, or resets its geometry, that pattern is more telling than any static frame. Watch the middle of every shot, not the beginning.


