发表机构
Gaoling School of Artificial Intelligence, Renmin University of China; Harvard University(中国人民大学中关村人工智能学院; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
综述人工智能生成视频检测,将任务重定义为事实保真度验证,提出视觉-语言双视角分类法,基于对221篇论文回顾,综合生成范式、审视检测方法、回顾评估指标,讨论挑战并指出未来方向。
AI 中文摘要
人工智能生成视频(AIGC-V)的逼真度不断提高,使传统以伪像为中心的检测方法不再适用,需要从低级检查转向高级语义验证。本文对AIGC-V检测进行全面综述,将该任务重新定义为事实保真度验证,即视频中描绘的事件、实体和物理过程是否与现实世界事实一致。为使该快速发展的领域系统化,提出视觉-语言双视角分类法,将现有方法分为四层。基于对221篇论文的系统回顾,综合AIGC-V生成范式,审视检测方法格局,回顾评估指标和基准。最后讨论当前挑战并指出未来方向。
英文摘要
The evolving realism of AI-generated Videos (AIGC-V) is rapidly rendering traditional artifact-centric detection insufficient, necessitating a paradigm shift from low-level inspection to high-level semantic verification. This paper presents a comprehensive survey of AIGC-V detection, reframing the task as Factual Fidelity Verification, which asks whether the events, entities, and physical processes depicted in a video are consistent with real-world facts. To systematize this rapidly evolving field, we propose a Vision-Language Dual-View taxonomy that organizes existing methods into a hierarchical, four-layer landscape, spanning intrinsic cue analysis, spatiotemporal consistency modeling, cross-modal consistency reasoning, and language-guided world-level reasoning. This dual-view framing highlights a fundamental transition from artifact matching in traditional deepfake detection to evidence-based semantic verification enabled by vision-language models and agentic reasoning pipelines. Based on a systematic review of 221 works, we synthesize AIGC-V generation paradigms, survey the landscape of detection methods, and review evaluation metrics and benchmarks in line with proposed views. Finally, we discuss current challenges and identify promising directions toward robust, explainable, and trustworthy detection.
Comments51 pages, accepted by ACL 2026
Journal refAssociation for Computational Linguistics 2026 pages 32221 to 32255
DOI:10.18653/v1/2026.findings-acl.1613