发表机构
Information Engineering University(信息工程大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI生成视频检测中因捷径学习导致跨域性能差的问题,提出G2VD框架,用反事实干预管道生成样本并对齐,设计因果解缠分类器,提升了跨域性能。
AI 中文摘要
人工智能生成视频的快速发展带来了越来越多的安全风险,需要强大的具有强跨域泛化能力的检测器。现有方法在域内评估中取得了有希望的结果,但在测试未见生成器时性能往往大幅下降。本文提出了G2VD,一个基于反事实干预和因果解缠的可泛化人工智能生成视频检测框架。首先,G2VD引入了一个反事实干预管道(CFIPipeline),通过变分自编码器(VAEs)生成受控的反事实样本,然后进行频域和像素域对齐,从而鼓励检测器专注于生成器内在线索。在此干预过程的基础上,我们进一步设计了一个因果解缠分类器,由两个具有不同分类目标的域锚定分支组成,并结合基于HSIC的独立性约束,以鼓励将任务相关线索与域特定偏差分离。在四个公共数据集上,G2VD显示出强大的平均跨域性能,并在匹配的主干上取得了一致的增益。在具有挑战性的GenVidBench跨域设置中,它超过了90%的准确率,AUC接近0.95。值得注意的是,此性能仅使用10%的原始训练数据即可获得。代码可在此https URL上获取。
英文摘要
Rapid advances in AI video generation pose increasing security risks and call for reliable detectors with strong cross-domain generalization. Although existing methods perform well under in-domain evaluation, their performance degrades substantially on unseen generators. A key reason is shortcut learning, where detectors rely on domain-specific bias rather than intrinsic forensic cues. To address this issue, we propose G2VD, a generalizable AI-generated video detection framework based on counterfactual intervention and causal disentanglement. First, G2VD introduces a counterfactual intervention pipeline (CFIPipeline) that constructs counterfactual samples through VAE-based reconstruction and subsequent frequency-domain and pixel-domain alignment, thereby weakening spurious correlations between domain-specific bias and authenticity labels. Building on this intervention, we further design a causal disentanglement classifier that combines two domain-anchored branches with complementary objectives and a constraint based on the Hilbert-Schmidt Independence Criterion (HSIC), encouraging the causal and non-causal representations to capture intrinsic forensic cues and domain-specific bias, respectively. Experiments across four public datasets demonstrate strong cross-domain performance and consistent gains over baseline methods. In the challenging GenVidBench setting, G2VD achieves over 90\% overall ACC, with improvements of 0.194 in F1 and 0.104 in AUC over comparable state-of-the-art methods, while using only 10\% of the available training data. Code is available at https://github.com/DMOSCAR-98/G2VD.