arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HumanForge:一个具有多智能体伪造原理的以人类为中心的深度伪造视频基准

HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-Agent Forgery Rationales

Wenbo Xu, Zhimin Chen, Xiaojie Liang, Hengrui Liu, Ziqi Sheng, Wei Lu

arXiv 2607.08705首次发表:更新:

AI 中文总结

针对视频伪造给数字内容取证带来的挑战,引入HumanForge数据集,提出基于LangGraph的Gen2Anno多智能体管道构建并注释数据集,经测试展现了零样本泛化和细粒度推理的重大挑战,代码和数据集将公开。

AI 中文摘要

视频扩散模型和时间编辑工具的快速发展使得生成高度逼真的以人为中心的视频成为可能,这给数字内容取证带来了前所未有的挑战。现有基准主要集中在面部交换或全局文本到视频合成,忽略了人与物体或人与人交互以及多模态对齐的关键维度。为解决这些限制,我们引入了HumanForge,一个统一、大规模且多范式的以人为中心的视频伪造数据集。为构建和注释该数据集,我们提出了Gen2Anno,一个基于LangGraph构建的模块化主动多智能体管道。Gen2Anno协调六个专门智能体,生成超过18K的高保真视频片段并产生结构化、对比性的全注释。使用先进传统探测器和大型多模态模型的广泛基准测试证明了在HumanForge上零样本泛化和细粒度推理的重大挑战。代码和数据集将公开发布。

英文摘要

Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, presenting unprecedented challenges to digital content forensics. Existing benchmarks primarily focus on face-swapping or global text-to-video synthesis, overlooking the crucial dimensions of multimodal alignment and complex human-object or human-human interactions. To address these limitations, we introduce HumanForge, a unified, large-scale, and multi-paradigm human-centric video forgery benchmark containing over 18,000 synthesized videos across four distinct scenarios: audio-driven, pose-driven, semantic-driven, and interaction. To construct and annotate this dataset without labor-intensive manual labeling or blind monolithic prompting, we propose Gen2Anno (Generation-to-Annotation), a cooperative multi-agent pipeline. Gen2Anno orchestrates six specialized agents-ranging from driving asset profiling to MoE-based reference analysis and closed-loop verification-to dynamically execute video synthesis and produce structured annotations containing binary authenticity labels, generative model attribution, and natural-language contrastive forgery rationales. By systematically contrasting expected states derived from generation provenance with actual visual observations, the framework generates logically grounded forensic reasoning chains. Extensive benchmarks using state-of-the-art traditional detectors and Vision-Language Models demonstrate the significant challenges of cross-generator generalization, perturbation robustness, and explainable reasoning on HumanForge. The code and dataset will be publicly released.

Comments18 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑