Guides
Sora, Veo, Kling: what AI video still gets wrong
Today's video generators are astonishing, but they share blind spots: physics, continuity, text and cause and effect. Where to look, model by model and in general.
Video Verifier team · · 3 min read
The leading video generators, including OpenAI's Sora, Google's Veo, Kuaishou's Kling, MiniMax's Hailuo, Luma's Ray and open models like Wan and Hunyuan, can now produce clips with realistic faces, natural light and even synchronised sound. Several of them generate speech and sound effects along with the picture.
But they all share the same basic design: they predict what a convincing video looks like, frame after frame. They don't simulate the world. That leaves fingerprints, and they show up in the same few places.
Cause and effect
Things in AI videos happen without a reason. A glass shatters before it hits the floor. A door opens before the hand reaches it. A ball changes direction in mid-air. Watch any moment where one thing should cause another and check that the order makes sense.
Object permanence
Generators have a short memory. When something leaves the frame or goes behind something else, it may come back different: a different number of buttons on a jacket, a bag that changed hands, a dog that changed breed. Pick an object and follow it.
Continuity across cuts
Many AI "news" videos and ads are stitched from separate clips. Between shots, a person's face, hairstyle, clothes or jewellery may not quite match. The background may change while the action supposedly continues.
Text, numbers and logos
Text is better than it was, but it remains unreliable: price tags, menus, street signs, jersey numbers and on-screen captions are often close to correct but wrong, or change between frames. Logos may be distorted versions of real brands.
Crowds and small faces
Main characters get the model's attention. Small faces in crowds, audiences and traffic are where you'll see people melting together, duplicate faces, or figures that freeze and then jump.
Hands, tools and contact
Hands are much better than a year ago, but contact is still hard: fingers gripping a steering wheel, a knife cutting food, a hand going into a pocket. Look at the exact moment of contact, frame by frame if you can.
Liquids, fire, smoke and cloth
These follow physical rules that are hard to fake for long. Water may pour upwards or vanish, smoke may freeze, and fabric may ripple in ways that don't match the wind or the body underneath.
Audio that's too good
Models that generate sound can produce voices that are perfectly clean even in a noisy street, footsteps that don't match the surface, or crowd noise that doesn't react to what's happening.
How detectors see what you can't
Every one of these slips is a disruption in how the video changes over time. That's why our in-house model focuses on motion: it measures how the content of each frame evolves through the clip, which catches many fakes that look perfect when paused. On our held-out tests it reached 95% balanced accuracy, though we're upfront that it's less sure on some of the newest generators, including Veo 3, Ray 2 and Hailuo 02, and on polished stock footage. More in how AI video detectors work.
Generators improve every few months. The specific glitches above will fade, but cause and effect, permanence and physics will stay hard for a long time. Train your eye on those.
Put it into practice
Next time a clip looks a little too perfect:
- Watch it once for the story, then once at a slower speed if you can.
- Follow one object, one hand and one background face.
- Check any text you can read.
- Run a free Quick Scan, or a Deep Scan if you're going to act on it.
For a fuller checklist, see 12 signs a video is AI-generated.