Measured 10 October 2026
How accurate is it? Every number we've measured.
Most AI detectors quote one big number. Here is everything we measured for our in-house video model, on videos it never saw while it was trained: AI tool by AI tool, including where it's weakest.
- of 601 AI videos from AI models it was never trained on were flagged
- 97%
- of those AI videos were called real by mistake
- ≤ 1%
- balanced accuracy on 11,952 held-out test videos from 20+ AI tools
- 95.3%
By AI video tool
Each clip was checked four times: as published, as a WhatsApp-quality copy, at 360p and heavily compressed. “Flagged” means our model leaned AI; “called real” means it was confident enough for a “Most likely real” answer, the mistake that matters most.
| AI video tool | Videos tested | Flagged as AI | Called real by mistake |
|---|---|---|---|
| OpenAI Sora516 videos | 516 | 97.5% | 0.8% |
| Kling (1.5, 1.6, 2.1)592 videos | 592 | 96.8% | 0.5% |
| Google Veo (Veo, 2, 3, 3.1)weaker216 videos | 216 | 88.9% | 2.3% |
| Runway (Gen-2, Gen-3)1,040 videos | 1,040 | 95.3% | 0.6% |
| Luma Dream Machine544 videos | 544 | 97.2% | 0.7% |
| Luma Ray 2weaker80 videos | 80 | 80.0% | 6.2% |
| MiniMax Hailuo (first version)328 videos | 328 | 95.1% | 1.5% |
| MiniMax Hailuo 02weaker92 videos | 92 | 84.8% | 5.4% |
| ByteDance (Seaweed, Seedance, PixelDance)932 videos | 932 | 95.3% | 1.4% |
| Pika (1.0, 2.2)160 videos | 160 | 91.9% | 0.0% |
| PixVerse (v2, v3)516 videos | 516 | 98.1% | 0.2% |
| Alibaba Wan (Wanx, 2.1)628 videos | 628 | 96.0% | 1.0% |
| Tencent HunyuanVideo372 videos | 372 | 91.7% | 1.3% |
| Zhipu CogVideoX / Qingying712 videos | 712 | 92.6% | 1.7% |
| Shengshu Vidu252 videos | 252 | 95.6% | 1.6% |
| Moonvalley Mareyweaker132 videos | 132 | 87.9% | 6.8% |
| Genmo Mochi 140 videos | 40 | 90.0% | 0.0% |
| Google VideoPoet368 videos | 368 | 97.0% | 0.3% |
| Stable Diffusion video and AnimateDiff920 videos | 920 | 99.6% | 0.0% |
| AI videos found on Wikimedia Commons (made with many tools)weaker60 videos | 60 | 58.3% | 10.0% |
Real videos
How often real footage was read as real (the other half of being accurate).
| Kind of real video | Videos tested | Read as real |
|---|---|---|
| Everyday real videos (Wikimedia Commons, hundreds of uploaders)3,320 videos | 3,320 | 96.5% |
| Polished stock footage (Pexels)weaker92 videos | 92 | 55.4% |
Where it's weakest, and what we do about it
- Polished stock footage often looks “too smooth” to the model, so real stock clips are read as real only about half the time. When the model isn't sure, the result says so instead of guessing.
- The newest AI tools (Luma Ray 2, Hailuo 02, Moonvalley Marey, Google Veo) are caught less often than older ones. We retrain as new tools appear.
- AI videos found in the wild, edited, re-encoded and mixed with real footage, are the hardest case for every detector, ours included (58% flagged on a small sample).
- Viral “breaking news” fakes of disasters and attacks are where detectors miss most. That's why a result is never “Most likely real” unless our video model is confident, and why links are also checked for AI labels in the post and for published fact-checks.
How we test
Held-out test. 2,988 clips (11,952 versions) set aside before training and never used to train or tune the model. Real and AI videos made from the same prompt stay on the same side of the split, so the model can't pass by recognising a scene. Sources, all under open licences: VideoGen-Eval (Yang et al., MIT), DeepAction (Bohacek & Farid, AI videos CC BY 4.0, real videos under the Pexels License), Rapidata's public text-to-video preference sets (Apache-2.0) and Wikimedia Commons (public domain, CC0 and CC BY files).
Independent test. 601 AI videos from two AI models outside the training data (NVIDIA Cosmos-Predict 2.5, a family the model never saw: 96.2% flagged; Wan 2.2 14B (image to video), earlier Wan versions were in training: 98.2% flagged) and 84 real videos from uploaders the model never saw; 96.4% of the real videos were read as real. AI videos: jnzhang/GeneratedVideos (CC BY 4.0).
These numbers are for the video model alone. A full result also weighs the frame check, the voice check, Content Credentials, metadata and, for links, the post itself; the final answer is an experimental estimate, never a certainty.
Got a video you're not sure about?
Check it free