New benchmarks and datasets advance deepfake detection for audio, image, and video
ByPulseAugur Editorial·[9 sources]·
Researchers have introduced several new datasets and benchmarks aimed at improving the detection of deepfakes across various media. Echoes focuses on music deepfakes, emphasizing semantic alignment and provider diversity to create more robust detection models. VendorBench-100 provides a unified framework for evaluating deepfake image detection across commercial APIs, vision-language models, and open-source detectors, highlighting performance differences and metric disagreements. HumanForge addresses deepfake video detection with a human-centric approach, using a multi-agent pipeline for annotation and focusing on human-object and human-human interactions. Additionally, XPlainVerse offers a large-scale benchmark for explainable deepfake detection, introducing new metrics to assess the fidelity of natural language explanations.
AI
IMPACT
These advancements in deepfake detection datasets and benchmarks are crucial for developing more robust and trustworthy AI systems, particularly in combating misinformation and ensuring digital content integrity.
RANK_REASON
Multiple research papers introducing new datasets and benchmarks for deepfake detection.
arXiv:2603.23667v2 Announce Type: replace-cross Abstract: We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions. Echoes comprises 4,468 tracks (131 hours of audio) spanning …
arXiv cs.AI
TIER_1English(EN)·Sharayu N. Deshmukh, Md Rashidunnabi, Nelton Tiago Gemo, Kurundkar G. D., Mahamune M. R., Nilesh K. Deshmukh·
arXiv:2607.06254v1 Announce Type: cross Abstract: Deepfake image detection is currently served by three fundamentally different paradigms: commercial APIs, zero-shot vision-language models (LLMs), and open-source detectors. Despite their widespread use, these paradigms are rarely…
Deepfake image detection is currently served by three fundamentally different paradigms: commercial APIs, zero-shot vision-language models (LLMs), and open-source detectors. Despite their widespread use, these paradigms are rarely evaluated under a common protocol, making direct …
arXiv cs.AI
TIER_1English(EN)·Kaliki V Srinanda, M Manvith Prabhu, Hemanth K Mogilipalem, Jayavarapu S Abhinai, Vaibhav Santhosh, Aryan Herur, Deepu Vijayasenan·
arXiv:2604.17376v2 Announce Type: replace-cross Abstract: In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the poor generalization capability of existing methods. In this paper, we use an ensemb…
arXiv:2607.08705v1 Announce Type: new Abstract: Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, posing unprecedented challenges to digital content forensics. Existing benchmarks primaril…
Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, posing unprecedented challenges to digital content forensics. Existing benchmarks primarily focus on either face-swapping or global text-t…
arXiv:2607.07216v1 Announce Type: new Abstract: Deepfake (DF) technology poses a significant threat to information integrity, driving the need for robust detection methods. Most DF detectors only consider predicting a binary label for whether the input is real or fake, lacking th…
Deepfake (DF) technology poses a significant threat to information integrity, driving the need for robust detection methods. Most DF detectors only consider predicting a binary label for whether the input is real or fake, lacking the justification required for real-world applicat…
arXiv cs.CV
TIER_1English(EN)·Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh, Muhammad Haris Khan, Jianfei Cai, Abhinav Dhall·
arXiv:2607.03562v1 Announce Type: new Abstract: As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limiting reliability and user trust. Existing benchmarks mainly evaluate classificat…