How it works
Perceptual straightening: the science behind our detector
Real videos trace straight paths inside a vision model; AI videos bend and jitter. The research (ReStraV, NeurIPS 2025) behind our detector, and our own results.
Video Verifier team · · 3 min read
Most AI video detectors look at one frame at a time. Our in-house model looks at something frame-by-frame checks can't see: the path a video takes through time. The idea comes from neuroscience, was turned into an AI-video detector by researchers in 2025, and is the core of how Video Verifier decides. Here's the science in plain language, with our own numbers, including where it fails.
The idea: our brains straighten natural video
In 2019, neuroscientists Olivier Hénaff, Robbe Goris and Eero Simoncelli published a finding in Nature Neuroscience: when you watch natural video, your visual system turns the frames into a representation where the sequence follows a straighter path than the raw pixels do. They called it perceptual straightening. A straight path is easy to predict, which may be why brains do it: knowing where something is heading lets you anticipate the next moment.
The key detail: this works for natural footage of the real world. Artificial sequences weren't straightened the same way.
From brains to AI video: ReStraV
In 2025 a team from Bielefeld University, Google DeepMind, the Honda Research Institute and Cold Spring Harbor Laboratory asked whether AI-generated video breaks this pattern. Their paper, AI-Generated Video Detection via Perceptual Straightening (Internò, Geirhos, Olhofer, Liu, Hammer and Klindt), was published at NeurIPS 2025.
They passed each frame through DINOv2, an open vision model that, like the visual system, turns images into compact descriptions. Then they measured two things about the path those descriptions take over time:
- Curvature: how sharply the path bends from one frame to the next.
- Step size: how far it jumps between frames.
Real videos produced straighter, steadier paths. AI videos bent and jittered more, even when every single frame looked perfect. A small classifier on these measurements reached 97.17% accuracy and 98.63% AUROC on the VidProM benchmark, according to the paper.
Why AI video bends
Video generators are trained to make each frame look right and to keep neighbouring frames roughly consistent. What they don't have is a world that obeys physics: light, mass, momentum and cameras that move through real space. Small inconsistencies (a texture that shifts, a hand that drifts, a background that breathes) are often too subtle for a person to notice, but they make the path through a vision model's representation less straight.
How we use it
Our model builds on that idea:
- It samples 64 frames in four short bursts across the first minute of a video, and runs each through DINOv2-small (Apache-2.0 licensed), about 12 billion calculations per frame.
- It measures the curvature and step sizes of the path, and also reads fine texture from DINOv2's intermediate layers, where generators leave traces.
- A small classifier, trained by us, turns those measurements into a score. We trained it on clips from many current generators (Veo, Sora, Kling, Seedance, Wan, Hunyuan, Pika, Ray, Hailuo and more) and on real footage from Wikimedia Commons uploaded by hundreds of different people, each clip in four qualities, including a WhatsApp-like copy.
It's one of eight checks. A large language model weighs all of them together before you see an answer.
Our results
- 95.3% balanced accuracy on 11,952 held-out test clip versions the model never saw in training.
- In an independent test (October 2026) it flagged 97.0% of 601 AI videos, including 96.2% from NVIDIA Cosmos, a family it never saw in training, and passed 96.4% of 84 real videos.
- Each clip was tested as published, as a WhatsApp-quality copy, at 360p and heavily compressed.
Full tables by AI tool are on our accuracy page.
Where it fails
We publish weak spots because a detector that hides them gets trusted when it shouldn't be:
- Polished stock footage: only 55.4% of Pexels-style stock clips passed as real. Our best guess is that stabilised, colour-graded and sometimes upscaled stock video shares traits with generated video.
- Some newer generators: Luma Ray 2 (80.0% flagged), Hailuo 02 (84.8%) and Google Veo (88.9%) are harder than average.
- Real footage with a false caption: the picture is real, so the model is right to say so. Context is a different question; see our 5-minute checklist.
When the signals don't agree, the answer says "not conclusive" instead of guessing.
Try it on a video you already know the answer to. Start a free scan, with no sign-up.