视频中可解释的深度伪造检测:显式取证特征与时序建模
Interpretable Deepfake Detection in Videos via Explicit Forensic Features and Temporal Modeling
浏览论文内容
中文总结 AI 辅助
本文提出一种可解释的视频深度伪造检测框架,通过显式编码多域取证特征并利用LSTM建模时序依赖,在四个基准上取得高F1分数,兼具鲁棒性与跨数据集泛化能力。
中文摘要 AI 辅助
视频中的深度伪造检测仍然具有挑战性,因为被操纵的内容可能在帧级别上看起来视觉一致,同时表现出细微的时序不一致性。本文提出了一种可解释的深度伪造检测框架,该框架对视频序列中面部特征的空间和时间一致性进行建模。与依赖隐式表示的端到端深度模型不同,所提出的方法显式地编码了基于物理的取证线索,从而实现透明的分析并提高跨数据集泛化能力。该流程将视频转换为身份一致的面部轨迹,将其分割为固定长度的时间窗口,并使用68个结构化描述符表示每一帧,这些描述符涵盖四个互补领域:光度、纹理、几何和基于压缩的特征。这些描述符提供了操作伪影的紧凑多域表示,并由长短期记忆(LSTM)网络处理,以捕获时间依赖性和细微的不规则性。在四个基准数据集(FaceForensics++、Celeb-DF v2、DeepFake检测挑战赛(DFDC)的精选子集以及DeeperForensics)上的评估分别产生了强大且一致的F1分数:98.0%、91.0%、97.6%和96.2%。该方法还展示了良好的跨数据集泛化能力,为视频深度伪造检测提供了一种稳健且可解释的解决方案。
英文摘要
Deepfake detection in videos remains challenging, as manipulated content may appear visually consistent at the frame level while exhibiting subtle temporal inconsistencies. This paper introduces an interpretable deepfake detection framework that models spatially and temporally coherent facial features in video sequences. Unlike end-to-end deep models relying on implicit representations, the proposed approach explicitly encodes physically grounded forensic cues, enabling transparent analysis and improved multi-dataset generalization. The pipeline transforms videos into identity-consistent facial trajectories, segments them into fixed-length temporal windows, and represents each frame using 68 structured descriptors spanning four complementary domains: photometric, textural, geometric, and compression-based features. These descriptors provide a compact multi-domain representation of manipulation artifacts and are processed by a Long Short-Term Memory (LSTM) network to capture temporal dependencies and subtle irregularities. Evaluation on four benchmark datasets, FaceForensics++, Celeb-DF v2, a curated subset of the DeepFake Detection Challenge (DFDC), and DeeperForensics, yields strong and consistent F1-scores of 98.0%, 91.0%, 97.6%, and 96.2%, respectively. The approach also demonstrated a good cross-dataset generalization, providing a robust and interpretable solution for video deepfake detection.
发表机构
- University of Quebec in Outaouais(魁北克大学乌塔韦校区)
机构由 AI 辅助整理,请以论文原文为准。