Researchers have introduced BabelFake, a new multilingual audio-visual DeepFake benchmark designed to address limitations in existing datasets. BabelFake includes 1,323 hours of footage from 496 individuals across five languages: English, German, Italian, French, and Spanish. The benchmark utilizes a modular pipeline that combines modern video manipulation techniques with voice cloning engines, differentiating between visual-only and joint audio-visual manipulations. Evaluations show that detection difficulty varies with the generation method, and performance degrades significantly when authentic audio is retained. Furthermore, human perception of realism does not always align with machine detection difficulty. AI
IMPACT This benchmark could improve the robustness of DeepFake detection models across various languages and manipulation techniques.
RANK_REASON The item describes a new benchmark dataset for audio-visual DeepFake detection, published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →