arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考读出方式:解锁用于人工智能生成视频检测的视频主干网络

Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection

Manni Cui, Ziheng Qin, ZiAn Wang, Ruiqi Liu, Dianyuan Zou, Jianglan Wei, Han Zhou, Yu Liu, Jingrui Xu, Wenhao Wang, Zhenyu Zhang

arXiv 2607.15321首次发表:更新:

发表机构

Huazhong University of Science and Technology; Institute of Automation, Chinese Academy of Sciences; Jilin University; Vast Intelligence Lab(华中科技大学; 中国科学院自动化研究所; 吉林大学; 旷视智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对AI生成视频检测中视频主干网络表现不佳的问题,提出速度门控补丁速度剖析(V-PVP)方法,通过替换聚合层激活冻结主干网络时间潜力,提升检测性能,在AIGVDBench上达95.28的AUC。

AI 中文摘要

人工智能生成的视频(AIGV)通常包含由帧间不一致而非单个帧内产生的细微时间伪影。因此,捕获此类伪影的检测器应受益于视频预训练主干网络而非仅图像预训练的探测器。然而,在实践中,具有标准全局读出的视频主干网络在AIGV基准测试中往往无法超越强大的图像预训练探测器。我们将这种差距归因于读出中的过度时空聚合。视频预训练主干网络倾向于将每个帧压缩成单个全局描述符。这种压缩抑制了局部补丁级别的时间动态并丢弃了补丁间关系,而这些正是AIGV检测最可靠依赖的线索。基于此,我们提出了速度门控补丁速度剖析(V-PVP),一种轻量级读出方式,仅用补丁速度场上的两个并行流替换聚合层,仅增加约0.5M可训练参数。V-PVP作为一个通用的即插即用模块,在端到端微调及线性探测设置下,能持续提升不同视频主干网络的性能。我们的方法在AIGVDBench上达到了95.28的AUC,同时保持主干网络完全冻结。结果表明,简单地替换聚合层就能重新激活冻结视频主干网络的时间潜力,恢复其在AIGV检测上的优势。代码可在该https链接获取。

英文摘要

AI-generated videos (AIGVs) typically contain subtle temporal artifacts that arise from inter-frame inconsistencies rather than within individual frames. A detector that captures such artifacts should therefore benefit from video pretrained backbones over image only ones. In practice, however, video backbones with standard global readouts often fail to outperform strong image pretrained probes on AIGV benchmarks. We attribute this gap to excessive spatiotemporal aggregation in the readout. Video pretrained backbones tend to compress each frame into a single global descriptor. This compression suppresses local patch level temporal dynamics and discards inter patch relations, which are precisely the cues that AIGV detection most reliably depends on. Based on this, we propose Velocity Gated Patch Velocity Profiling (V-PVP), a lightweight readout that replaces only the aggregation layer with two parallel streams over the patch velocity field, adding only about $0.5$M trainable parameters. V-PVP serves as a general plug-and-play module that consistently improves performance across diverse video backbones under both end-to-end fine-tuning and linear probing settings. Our method reaches \textbf{95.28} AUC on AIGVDBench while keeping the backbone fully frozen. The results show that simply replacing the aggregation layer reactivates the temporal potential of frozen video backbones, restoring their advantage on AIGV detection. Code is available at https://anonymous.4open.science/r/PVP-81B3/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑