发表机构
Digital Future Lab, Flanders Make, Hasselt University(数字未来实验室,弗兰德斯制造,哈瑟尔特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示了四种基于运动的AI生成视频检测器存在显著预处理和采样偏差,导致性能在无运动偏差的数据集上降至随机水平,而频率方法更具泛化性。
AI 中文摘要
近年来,AI生成视频的视觉质量大幅提升,使得人类越来越难以区分真实与合成媒体。本文评估了四种最先进的基于运动的AI生成视频检测器的鲁棒性和适用性。我们发现了这些方法中存在显著的预处理和采样偏差,并证明这些偏差对其报告性能有实质性贡献。此外,我们发现这些检测器对其评估数据集特有的运动模式高度敏感,其中AI生成视频通常比真实视频表现出更少的帧间运动。我们表明,当在不存在这种运动偏差的数据集上评估时,所有检测器的性能都降至接近随机水平。此外,通过数据集重平衡和简单的空间增强,我们观察到所有评估模型的性能严重下降。相比之下,我们发现一种现有的基于频率的检测器在所有评估数据集上保持强劲性能,表明基于频率的方法可能为AI生成视频检测提供一条更具泛化性的路径。我们希望我们的工作能提高对这些漏洞的认识,并鼓励开发更具代表性、无偏差的数据集和更鲁棒的评估协议。
英文摘要
The visual quality of AI-generated videos has improved drastically in recent years, making it increasingly difficult for humans to distinguish between real and synthetic media. In this work, we evaluate the robustness and applicability of four state-of-the-art motion-based AI-generated video detectors. We identify significant preprocessing and sampling biases in three of the four methods and demonstrate that they account for a substantial portion of their reported performance. Furthermore, we find that these detectors are highly sensitive to motion patterns specific to their evaluation datasets, where AI-generated videos generally exhibit less inter-frame movement than real videos. We show that for all detectors, performance collapses to near-random levels when evaluated on a dataset that does not contain this motion bias. Additionally, through dataset rebalancing and the application of simple spatial augmentations, we observe severe performance degradation across all evaluated models. In contrast, we find that an existing frequency-based detector maintains strong performance across all evaluated datasets, suggesting that frequency-based approaches may offer a more generalizable path forward for AI-generated video detection. We hope that our work raises awareness towards these vulnerabilities and encourages the development of more representative, unbiased datasets and more robust evaluation protocols.