ManiVid:统一且可解释的篡改视频取证分析
ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos
- Shanghai Jiao Tong University(上海交通大学)
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Fudan University(复旦大学)
- The Chinese University of Hong Kong(香港中文大学)
- Shanghai AI Laboratory(上海人工智能实验室)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对篡改视频取证中数据稀缺和MLLM难以利用低级线索的问题,提出ManiVid任务、ManiVid-38K数据集及ManiVidLens统一框架,实现检测、定位与解释,显著提升定位和解释性能。
AI中文摘要:
AI生成视频(AIGV)的快速发展增加了欺骗性视频篡改所带来的风险。与完全合成的视频不同,篡改视频保留了大部分源内容,仅修改局部区域,这使得取证分析尤其具有挑战性。现有的视频伪造研究在数据和方法论两方面均面临两个局限:(1)针对篡改视频的高质量数据集和基准仍然稀缺。(2)多模态大语言模型(MLLMs)将伪造分析扩展到二元分类之外,但难以利用低级取证线索并提供精确的像素级定位。具体而言,我们引入了ManiVid,一个统一的取证分析任务,涵盖篡改视频的伪造检测、伪影定位和异常解释。我们构建了ManiVid-38K,这是首个将成对的、开放词汇的通用视频局部篡改与真实性标签、伪造掩码和异常解释相结合的数据集。它包含约19,000对人工验证的真实-伪造视频对,大部分为1080P分辨率,在2种范式下使用15种强大的生成模型生成。我们从中采样1,000对用于ManiVidBench,在六种篡改类型和生成模型之间保持平衡,以确保公平评估。我们进一步提出了ManiVidLens,一个用于可解释视频伪造分析的统一框架。其取证证据路由器(Forensic Evidence Router)为多模态推理和视频分割提供共享的低级取证证据。其提示蒸馏模块(Prompt Distill Module)将定位状态转换为语义和几何提示,并蒸馏空间先验用于掩码解码和全视频传播。ManiVidLens在伪影定位(+21.1% mIoU;+21.3% J&F)和异常解释(+131.3% ROUGE-L;+9.9% CSS)方面相较于最强对比方法取得了相对提升。其伪造检测性能与专用分类器相当(0.914准确率;0.913 F1分数)。
英文摘要:
Rapid advances in AI-generated video (AIGV) have increased the risks posed by deceptive video manipulation. Unlike fully synthetic videos, manipulated videos retain most source content and alter only localized regions, making forensic analysis particularly challenging. Existing video forgery research faces two limitations in both data and methodology: (1) High-quality datasets and benchmarks tailored for manipulated videos remain scarce. (2) Multimodal large language models (MLLMs) extend forgery analysis beyond binary classification but struggle to use low-level forensic cues and provide precise pixel-level grounding. Specifically, we introduce ManiVid, a unified forensic analysis task covering forgery detection, artifact grounding, and anomaly explanation for manipulated videos. We construct ManiVid-38K, the first dataset to combine paired, open-vocabulary localized manipulations of general videos with authenticity labels, forgery masks, and anomaly explanations. It comprises about 19K manually verified real-fake video pairs, mostly at 1080P resolution, generated under 2 paradigms with 15 powerful generation models. We sample 1K pairs for ManiVidBench, balanced across six manipulation types and generation models for fair evaluation. We further propose ManiVidLens, a unified framework for explainable video forgery analysis. Its Forensic Evidence Router supplies shared low-level forensic evidence for multimodal reasoning and video segmentation. Its Prompt Distill Module converts grounding states into semantic and geometric prompts and distills spatial priors for mask decoding and full-video propagation. ManiVidLens achieves relative gains over the strongest comparison methods in artifact grounding (+21.1% mIoU; +21.3% J&F) and anomaly explanation (+131.3% ROUGE-L; +9.9% CSS). Its forgery detection remains comparable to dedicated classifiers (0.914 Acc; 0.913 F1).