基于首位数字梯度统计的合成视频检测通用可解释框架
A Generalizable and Explainable Framework for Synthetic Video Detection Using First-Digit Gradient Statistics
浏览论文内容
中文总结 AI 辅助
提出基于Sobel梯度首位数字统计的通用可解释AI视频检测框架,结合线性判别分析与多层感知器,在多个基准上验证,并揭示零样本检测的失败现象。
中文摘要 AI 辅助
AI视频生成器不仅越来越难以检测,而且被用于生成从风景到街景再到动物视频的多样化场景。这造成了一个问题:基于CNN的检测器虽然有效,但无法洞察其内部工作机制,而基于取证学的检测器通常针对特定场景预训练,或变得过于复杂而难以得出有意义的见解。我们提出了一种新颖的AI视频检测方法,使用Sobel梯度值并结合首位数字定律进行分析。通过线性判别分析,我们可视化判别信号,同时使用多层感知器进行分类。该检测方法没有生成器或场景特定的特征,模型也不了解容器格式、编解码器、比特率或压缩伪影。模型在GenBuster-200K、GenBusterBench、GenVA、FaceForensics++ C23和CelebDF上进行了训练和测试。我们还展示了即使特征集携带判别信号,零样本检测也会失败。
英文摘要
AI video generators have not only become harder to detect but are used to generate a diverse set of scenarios from landscapes to street views to animal videos. This creates a problem where CNN-based detectors are effective but offer no insight into their inner workings, while forensics-based detectors are often pretrained for a set scenario or become too complex to derive meaningful insights. We present a novel approach to AI video detection using Sobel gradient values analysed with the first-digit law. Using linear discriminant analysis, we visualise the discriminatory signal, while a multi-layer perceptron is used for classification. The detection method has no generator- or scenespecific features, and the model has no knowledge of container formats, codec, bitrate, or compression artefacts. The model is trained and tested on GenBuster-200K, GenBusterBench, GenVA, FaceForensics++ C23, and CelebDF. We also show how zero-shot detection fails even though the feature set carries a discriminatory signal.
发表机构
- IIT Jodhpur(印度理工学院焦特布尔分校)
- GGSIPU(古鲁·戈宾德·辛格·因德拉普拉萨塔大学)
- MITS Gwalior(瓜廖尔米拉理工学院)
机构由 AI 辅助整理,请以论文原文为准。