发表机构
City University of New York(纽约城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对资源受限环境下部分篡改AI视频检测,提出轻量级全帧检测器,通过知识蒸馏与多种正则化技术,在边缘设备上实现高效检测,显著缩小与教师模型的性能差距。
AI 中文摘要
生成式视频模型的普及已将实际检测威胁从完全伪造的片段转移到部分篡改的片段。尽管现代检测器使用4亿以上参数的基础骨干网络实现了较强的准确性,但其资源占用阻碍了边缘部署。在本文中,我们提出了一种轻量级全帧检测器,用于检测部分篡改的AI生成视频,设计用于在无需人脸检测预处理的边缘硬件上部署。该系统通过结合温度退火软标签迁移、注意力多样性正则化、帧级监督以及残差特征适配器的流水线,将DINOv2-Base教师模型蒸馏到冻结的MobileNetV3-Small学生模型中,其中残差特征适配器对ImageNet特征进行条件化以进行伪影检测。我们还针对部分篡改场景特有的两种失败模式:对合法场景切换的误报,通过视频内时间硬负样本解决;以及对占主导地位的纯真实类别的阈值级校准错误,通过校准感知采样解决。在包含55,393个样本的拼接测试集上,伪造帧比例从6.2%到31.2%的评估表明,学生模型缩小了与DINOv2-Base教师模型(AUC 0.766)之间58%的差距,同时在RTX A4000上以每16帧片段3.65毫秒的速度运行,且检查点大小为150.4 MB,符合边缘设备的内存和延迟预算。
英文摘要
The proliferation of generative video models has shifted the practical detection threat from fully fabricated clips to partially manipulated footages. Although modern detectors achieve strong accuracy using foundation backbones of 400M+ parameters, their resource footprint precludes edge deployment. In this paper, we present a lightweight full-frame detector for partially manipulated AI-generated video, designed for deployment on edge hardware without face-detection preprocessing. The system distills a DINOv2-Base teacher into a frozen MobileNetV3-Small student through a pipeline that combines temperature-annealed soft-label transfer, attention-diversity regularization, frame-level supervision, and a residual feature adapter that conditions ImageNet features for artifact detection. We additionally target two failure modes specific to the partial-manipulation regime: false positives on legitimate scene cuts, addressed through within-video temporal hard negatives; and threshold-level miscalibration on the dominant pure-real class, addressed through calibration-aware sampling. Evaluation on a 55,393-sample spliced test set across fake-frame ratios from 6.2% to 31.2% demonstrates the student model closing 58% of the gap to the DINOv2-Base teacher (AUC 0.766) while running at 3.65 ms per 16-frame clip on RTX A4000 with a 150.4 MB checkpoint compatible with edge-device memory and latency budgets.
Comments7 Pages