发表机构
Chung-Ang University; GS. of AI, Chung-Ang University; Dept. of Advanced Imaging, GSAIM, Chung-Ang University; Dept. of Metaverse Convergence, Chung-Ang University(中央大学; 中央大学人工智能学院; 中央大学先进影像系(GSAIM); 中央大学元宇宙融合系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有深度伪造检测方法无法适配视频特有线索的问题,提出MSFD框架,通过频域分解与跨模态去相关损失实现各模态知识保留,在持续视频深度伪造检测中性能优于现有最优方法。
AI 中文摘要
高质量视频深度伪造的持续出现,要求检测器能持续适配新的伪造模式,但现有针对深度伪造图像设计的方法无法捕捉视频特有的线索。与仅包含空间伪影的深度伪造图像不同,深度伪造视频会在空间和时间轴上留下明显证据,因此在序列模型更新期间需分别保留各模态信息。为克服这一局限,我们提出一种持续视频深度伪造检测框架——模态特定频率蒸馏(Modality-Specific Frequency Distillation, MSFD),该框架在频域中将视频特征显式分解为空间、时间和时空模态。这种分解可实现各模态的独立保留,因为不同类型的深度伪造视频在不同任务中对空间和时间线索的依赖程度各异。此外,MSFD采用跨模态去相关损失,鼓励时空表征与单模态线索保持正交。大量实验表明,在各类持续视频深度伪造场景中,我们的框架相比现有最优方法实现了更强的适应性,且能更有效地保留性能。
英文摘要
The continuous emergence of high-quality video deepfakes requires detectors that continually adapt to new forgery patterns, yet existing approaches, which are designed for deepfake images, fail to capture video-specific cues. Unlike deepfake images that contain only spatial artifacts, deepfake videos leave distinct evidence along both spatial and temporal axes, necessitating the separate preservation of each modality during sequential model updates. To overcome this limitation, we introduce a continual deepfake video detection framework, Modality-Specific Frequency Distillation (MSFD), that explicitly decomposes video features into spatial, temporal, and spatiotemporal modalities in the frequency domain. This decomposition enables independent preservation of each modality, as different deepfake video types exhibit varying reliance on spatial and temporal cues across tasks. Furthermore, MSFD adopts a cross-modality decorrelation loss that encourages spatiotemporal representations to remain orthogonal to single-modality cues. Extensive experiments show that our framework achieves stronger adaptation and preserves performance more effectively than state-of-the-art methods across diverse continual deepfake video scenarios.
CommentsAccepted by ECCV 2025. Code will be available at github.com/rama0126/MSFD