发表机构
The Hong Kong Polytechnic University; Carnegie Mellon University(香港理工大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI生成视频检测泛化性差的问题,提出MTOR方法,融合视觉与字幕文本多模态语义,并利用时间过度规律性特征,在多个基准上超越现有方法。
AI 中文摘要
视频生成技术的快速发展缩小了真实视频与合成视频之间的感知差距,使得可泛化的AI生成视频检测日益具有挑战性。现有检测器主要依赖视觉表示,而忽略了由字幕导出的文本语义。同时,细粒度视觉表示中的时间规律性也受到有限关注。我们发现,由字幕导出的文本表示能为全局视觉表示提供互补的判别线索。进一步分析揭示,AI生成视频表现出更强的时间持续性和更低的时间变异性,我们将这种模式称为时间过度规律性(TOR)。基于这些发现,我们提出了MTOR,包含一个多模态分支和一个TOR组件。多模态分支整合了全局视觉表示和由字幕导出的文本表示,而TOR组件在三个层面建模时间过度规律性:粗粒度的帧间连续性、细粒度的令牌对应性以及帧到视频的稳定性。在覆盖46个生成器变体的五个基准上的广泛评估表明,与16个代表性基线相比,我们的方法达到了最先进的整体性能,而鲁棒性实验确认了其对十二种真实世界视频扰动的强韧性。代码和模型将在此https网址发布。
英文摘要
The rapid evolution of video generation has narrowed the perceptual gap between authentic and synthetic videos, making generalizable AI-generated video detection increasingly challenging. Existing detectors predominantly rely on visual representations, leaving caption-derived textual semantics underexplored. Meanwhile, temporal regularity in fine-grained visual representations has received limited attention. We find that caption-derived textual representations provide complementary discriminative cues to global visual representations. Our analysis further reveals that AI-generated videos exhibit stronger temporal persistence and lower temporal variability, a pattern we term temporal over-regularity (TOR). Based on these findings, we propose MTOR with a multimodal branch and a TOR component. The multimodal branch integrates global visual and caption-derived textual representations, while the TOR component models temporal over-regularity at three levels: coarse inter-frame continuity, fine-grained token correspondence, and frame-to-video stability. Extensive evaluations on five benchmarks covering 46 generator variants demonstrate state-of-the-art overall performance against 16 representative baselines, while robustness experiments confirm strong resilience to twelve real-world video perturbations. Code and models will be released at https://github.com/hwang-cs-ime/MTOR.
Comments18 pages, 4 figures, 19 tables