How it works
How AI video detectors work, and why none is 100% sure
Frame scanners, motion models, voice checks and Content Credentials: what each one looks for, where each one fails, and why a good detector shows its evidence.
Video Verifier team · · 4 min read
"Is this video AI?" sounds like a yes-or-no question. In practice, a detector is a set of separate checks, each good at spotting some fakes and blind to others. Knowing what each check does makes it much easier to judge a result, ours included.
Check 1: Looking inside each frame
The oldest approach treats a video as a stack of images. A model trained on millions of real photos and AI images learns the statistical fingerprints that image generators leave in pixels: unusual textures, colour patterns and frequency artefacts that people can't see.
Strength: good at catching images and video from generators similar to the ones it was trained on.
Weakness: it looks at frames one at a time, so it can't see how things move. In our tests, a well-known open frame detector missed many clips from newer video models like Veo 3 and Hunyuan. Heavy compression also smooths away the fine traces it relies on.
Check 2: Looking at how the picture moves
Video generators are judged frame by frame on how good each picture looks, but motion is where they still slip. Real cameras record a world that obeys physics, so the content of real footage tends to change smoothly and consistently from one frame to the next.
Our in-house model is built on this idea. It turns each frame into a compact description using DINOv2, a widely used open vision model, then measures how that description travels through time: how straight, smooth and consistent the path is. A classifier trained on thousands of real and AI clips learns which motion patterns belong to cameras and which belong to generators.
We trained it on clips from many current generators, including Veo, Sora, Kling, Seedance, Wan, Hunyuan, Pika, Ray and Hailuo, alongside real footage from thousands of different people, each in several qualities so it also learns what compression looks like. On 11,952 test clips it never saw during training, it reached 95% balanced accuracy. In a head-to-head on the same 400 held-out clips, it scored 97% where the frame-only detector scored about 70%.
Weakness: we're open about this one. It's less reliable on polished, stock-style real footage, which can look "too smooth", and on a few of the newest generators. That's why it's one check among several, not the whole answer.
Check 3: Listening to the voice
Voice checks look for patterns that speech synthesis and voice cloning leave in audio. They help with videos where real footage has been given fake words. Honestly, this is the hardest area today: in our tests, open voice detectors scored only about 48 to 63% on modern voice clones. A voice result on its own should never decide anything. More on this in our guide to voice cloning scams.
Check 4: Reading Content Credentials
Some AI tools and cameras now sign their files with Content Credentials (the C2PA standard): a tamper-evident record of how the file was made. If a video carries a valid credential saying it was generated by Sora, that's the strongest evidence there is. The catch: most videos don't carry one, and re-uploading or screen-recording strips it. Absence proves nothing. See Content Credentials explained.
Check 5: Metadata and file forensics
A file's metadata can show the app that saved it, missing camera details, or audio and video tracks that don't line up. These are hints rather than proof, because social platforms rewrite metadata on upload, but they add context to the other checks.
Why no detector can be 100% sure
- New generators keep arriving. A model can only recognise patterns similar to what it has learned.
- Compression destroys evidence. Every re-upload and every WhatsApp forward throws away fine detail. Read more in why re-uploads make fakes harder to catch.
- Real and fake get mixed. A real video with an AI voice, or one AI-edited face in real footage, is neither fully real nor fully fake.
- Short clips carry less signal. Three seconds of video gives every check less to work with.
What a trustworthy result looks like
A single percentage with no explanation isn't enough to act on. A trustworthy detector should:
- Show which checks found what, and where in the video.
- Say plainly when a check was unavailable, instead of quietly scoring it.
- Say "too close to call" when the evidence is split, rather than guessing.
- Separate fully AI-made from real footage with AI added.
That's how Video Verifier reports every result: a headline like "92% AI" or "88% Real", a plain-language answer, and the evidence behind it. It's a strong signal for deciding what to trust, and it's built to be checked, not just believed.