AI 中文总结
针对深度伪造检测的现有方法泛化性不足问题,本文构建含10万视频的FaceVid-Forensics-100K数据集,提出多智能体取证推理框架,在域外测试集上性能优于GPT、Gemini等模型。
AI 中文摘要
生成式人工智能被恶意用于创建高度逼真的深度伪造视频,引发了严重的伦理问题,并对AI安全构成重大挑战。然而,现有的深度伪造视频基准对近期合成方法的覆盖有限,且通常缺乏可靠的细粒度文本标注。与此同时,传统检测器和多模态大语言模型(MLLMs),无论是作为单一模型运行还是依赖单一分析视角,往往无法捕捉到细微的伪造痕迹,限制了它们对新兴AI生成方法的泛化能力。为解决这些局限性,我们推出FaceVid-Forensics-100K,这是一个大规模深度伪造视频数据集,包含100,000个视频,涵盖换脸、人脸重演和全脸合成等33种合成方法,包括Seedance 2.0等近期生成器。该数据集提供视觉观察的细粒度文本标注和与判决一致的取证解释,这些解释由先进MLLMs驱动的多模型聚合与冲突解决流水线自动合成。基于此基准,我们提出一种多智能体取证推理框架,该框架采用四个专门的领域专家智能体,从纹理、光照、运动和物理四个视角独立分析伪造线索,随后由一个裁判智能体协调它们的报告,以生成最终预测及解释。对域外测试集的广泛评估表明,尽管该框架完全由小型开源MLLMs构成,但它在所有报告的指标上均优于包括闭源GPT和Gemini模型在内的所有方法,并在该基准上排名第一。项目页面可访问此https URL。
英文摘要
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine-grained textual annotations. Meanwhile, conventional detectors and multimodal large language models (MLLMs), whether operating as a single model or relying on a single analytical perspective, often fail to capture subtle forgery artifacts, limiting their generalization to emerging AI-generated methods. To address these limitations, we introduce FaceVid-Forensics-100K, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, including recent generators such as Seedance 2.0. The dataset provides fine-grained textual annotations of visual observations and verdict-consistent forensic explanations, automatically synthesized through a multi-model aggregation and conflict-resolution pipeline powered by advanced MLLMs. Building on this benchmark, we propose a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports to produce a final prediction together with an explanation. Extensive evaluations on out-of-domain test sets show that, despite being composed entirely of small open-source MLLMs, our framework outperforms all methods including closed-source GPT and Gemini models and ranks first across all reported metrics on this benchmark. The project page is available at https://xavierjiezou.github.io/ARGUS/.
Comments22 pages, 8 figures, 14 tables