超越模糊视觉线索:研究深度伪造视频中的生理扰动与跨模态不一致性
Beyond Ambiguous Visual Cues: Studying Physiological Disruptions and Cross-Modal Inconsistencies in Deepfake Videos
查看机构详情
- LIASD Laboratory(LIASD实验室)
- Shanghai Jiaotong University(上海交通大学)
- University of Paris 8(巴黎第八大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文构建高保真深度伪造数据集,提出双向协同注意力融合检测器,联合建模rPPG与面部行为,在多个数据集上超越现有方法,验证了生理与行为跨模态一致性在伪造检测中的有效性。
中文摘要 AI 辅助
最近的深度伪造检测研究日益将远程光电容积描记(rPPG)信号视为真实性线索。然而,现有基准缺乏生理真值,当前检测器也未能充分探索面部特征与生理动态之间的跨层级关系,通常依赖晚期融合或仅使用rPPG特征。在本文中,我们在已建立的真实rPPG数据集(COHFACE和UBFC-rPPG)上构建高保真深度伪造操作,以研究伪造如何同时扰乱自然生理信号和面部行为。基于此分析,我们提出了一种双向协同注意力融合检测器,联合建模rPPG和面部行为令牌。该机制显式捕获脉搏动态与面部运动之间的跨层级依赖,以学习鲁棒的联合真实性表示。使用主题不相交的5折评估进行的大量实验证明了我们方法的优越性。在构建的数据集上,使用换脸方法达到92.80%的AUC,使用运动迁移方法达到96.78%的AUC,我们的模型优于仅rPPG的单模态基线和最佳特征级融合方法。此外,在保持两个特征提取器固定的情况下,融合检测器在Celeb-DF-v2上的迁移学习结果达到91.20%的准确率和86.08%的AUC,这表明了其在目标域适应下的适用性。
英文摘要
Recent deepfake detection studies increasingly suggest remote photoplethysmography (rPPG) signals as an authenticity cue. However, existing benchmarks lack physiological ground truth, and current detectors underexplore the cross-level relationship between facial features and physiological dynamics, often relying on late fusion or rPPG features alone. In this paper, we construct high-fidelity deepfake manipulations on established real rPPG datasets (COHFACE and UBFC-rPPG) to investigate how forgeries disrupt natural physiological signals and facial behavior at the same time. Building on this analysis, we propose a bidirectional co-attention fusion detector that jointly models rPPG and facial behavior tokens. This mechanism explicitly captures the cross-level dependencies between pulse dynamics and facial motion to learn a robust, joint authenticity representation. Extensive experiments using a subject-disjoint 5-fold evaluation demonstrate the superiority of our approach. Achieving a 92.80\% AUC on constructed datasets using face swapping and 96.78\% AUC on motion transfer, our model outperforms both the rPPG-only single modality baseline and the best feature-level fusion methods. Furthermore, transfer-learning result of the fusion detector on Celeb-DF-v2 while keeping both feature extractors fixed achieves 91.20\% accuracy and 86.08\% AUC, which suggests applicability under target-domain adaptation.