DF26: We Cannot Tell Fake From Real Anymore
DF26 benchmark shows humans and state-of-the-art deepfake detectors perform near chance on videos generated by seven modern text-to-video models.
Researchers introduce DF26, a benchmark of 271 real and 2,420 fully synthetic videos created by seven modern video generation models, all depicting single-person public-speaking scenarios such as direct-to-camera recordings, official statements, and studio interviews. Human viewers and state-of-the-art deepfake detectors scored close to random chance at distinguishing fakes from real footage. The authors argue current evaluation protocols are insufficient and call for benchmarks that explicitly measure robustness to modern generative model distribution shifts.
- Benchmark spans 271 real and 2,420 synthetic videos from seven video models.
- Human and detector accuracy on DF26 is close to random chance.
- Focus on single-person public-speaking scenarios across studio and direct-to-camera settings.
- Authors call for evaluations robust to modern generative distribution shifts.
Full article96 words · extracted from huggingface.co · click to collapse
We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to modern generative model distribution shifts.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.07369