arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25775cs.CV

TRACE:面向AI生成视频检测的轨迹表示与一致性估计

TRACE: Trajectory Representation and Consistency Estimation for AI-Generated Video Detection

  • Zhejiang University(浙江大学)
  • University of the Chinese Academy of Sciences(中国科学院大学)
  • Nanjing University of Information Science and Technology(南京信息工程大学)
  • Zhejiang Gongshang University(浙江工商大学)

机构由 AI 辅助整理,请以论文原文为准。

Huangsen Cao, Hongkang chu, Siyao Yu, Xin Ding, Jianfeng Dong, Yongwei Wang

AI总结:

针对AI生成视频检测泛化性差的问题,提出TRACE框架,利用预训练视频DiT的速度响应作为可迁移取证线索,结合跨帧一致性建模与真实中心轨迹优化,在AIGVDBench上显著超越现有方法。

AI中文摘要:

近期生成式视频模型的进展使得合成视觉逼真内容成为可能,这给合成视频检测带来了重大挑战。现有检测器往往依赖于外观伪影、语义不一致性以及可能特定于生成器的时间模式,从而限制了其对未见合成模型的泛化能力。我们探究了预训练生成模型的响应是否能提供更具迁移性的取证线索。我们的关键观察是,在预训练的Flow Matching视频模型下,真实视频与AI生成视频表现出截然不同的速度响应。当使用不同的预训练视频生成骨干网络作为探针时,这种差异依然存在,表明速度响应提供了超越视觉伪影的可迁移取证信号。受此观察启发,我们提出了TRACE(轨迹表示与一致性估计),一种面向AI生成视频检测的生成过程感知框架。TRACE利用预训练的视频DiT作为速度场探针,在多个流时间点提取表示,并通过相邻帧之间的速度差建模跨帧一致性。我们进一步引入了一种真实中心轨迹优化目标,以鼓励生成器不变表示学习。在AIGVDBench上的大量实验表明,TRACE能够在不同生成器间有效泛化,在未见过的开源和闭源视频生成模型上大幅超越先前的最先进方法。

英文摘要:

Recent advances in generative video models have enabled the synthesis of visually realistic content, posing significant challenges to synthetic video detection. Existing detectors often rely on appearance artifacts, semantic inconsistencies, and temporal patterns that may be generator-specific, limitating generalization to unseen synthesis models. We investigate whether responses to a pretrained generative model provide more transferable forensic cues. Our key observation is that real and AI-generated videos exhibit distinct \emph{velocity responses} under a pretrained Flow Matching video model. This distinction persists when different pretrained video-generation backbones are used as probes, suggesting that velocity responses offer transferable forensic signals beyond visual artificts. Motivated by this observation, we propose \textbf{TRACE} (\emph{\underline{T}rajectory \underline{R}epresentation \underline{a}nd \underline{C}onsistency \underline{E}stimation}), a generation-process-aware framework for AI-generated video detection. TRACE leverages a pretrained video DiT as a velocity-field probe to extract representations at multiple flow time points, and models cross-frame consistency through velocity differences between adjacent frames. We further introduce a \emph{Real-Centered Trajectory Optimization} objective that encourages generator-invariant representation learning. Extensive experiments on AIGVDBench demonstrate that TRACE generalizes effectively across diverse generators, substantially outperforming prior state-of-the-art methods on unseen open- and closed-source video generation models.

↑