arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BabelFake:一个多语言音视频深度伪造基准

BabelFake: A Multilingual Audio-Visual DeepFake Benchmark

Carlotta Segna, Joel Tschesche, Anna Rohrbach

arXiv 2610.06339首次发表:更新:

发表机构

TU Darmstadt(达姆施塔特工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BabelFake是一个多语言音视频深度伪造基准,包含399k个片段和五种语言,通过模块化管道生成,用于评估检测器性能,发现检测难度受生成配对影响且人类感知与机器检测不一致。

AI 中文摘要

可靠且实用的音视频深度伪造检测需要反映多样化语言背景和现代数据合成管道的基准,涵盖视觉和音频操纵。然而,现有数据集主要包含英语说话者的视频,通常包含过时的操纵类型,或忽略音频模态。此外,许多数据集包含未同意被用于深度伪造创建的个人。我们引入了BabelFake,一个由同意参与的参与者录制的多语言音视频深度伪造基准。BabelFake包含来自496名个人的399k个片段(1,323小时),涵盖五种语言(英语、德语、意大利语、法语、西班牙语)。我们的模块化数据生成管道将11种现代视频操纵方法与4种语音克隆引擎配对,区分仅视觉(换脸)和联合音视频操纵(唇形同步和肖像动画)。通过基准测试最先进的检测器,我们表明检测难度取决于音视频生成配对,当保留真实音频时,性能显著下降。跨语言/人口统计评估揭示了不同检测器架构和训练数据之间的敏感性差异,而人类评估表明感知真实性和机器检测难度不一定一致。

英文摘要

Reliable and practical audio-visual DeepFake detection requires benchmarks that reflect diverse linguistic contexts and modern data synthesis pipelines for visual as well as audio manipulations. However, existing datasets predominantly contain footage of English-speakers, often include outdated manipulation types, or overlook the audio modality. Further, many datasets feature individuals who did not consent to be used in DeepFake creation. We introduce BabelFake, a multilingual audio-visual DeepFake benchmark recorded with consenting participants. BabelFake contains 399k clips (1,323 hours) from 496 individuals spanning five languages (English, German, Italian, French, Spanish). Our modular data generation pipeline pairs 11 modern video manipulation methods with 4 voice cloning engines, distinguishing visual-only (face swapping) and joint audio-visual manipulations (lip synchronization and portrait animation). By benchmarking state-of-the-art detectors, we show that detection difficulty depends on the audio-visual generation pairing, with substantial performance degradation when authentic audio is preserved. Cross-language/demographic evaluation reveals sensitivity varying across detector architectures and training data, while human evaluation reveals that perceived realism and machine-detection difficulty do not necessarily align.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑