arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

互补的rPPG导出与唇部区域频率线索用于说话人脸深度伪造检测

Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection

Othmane Harraq, Tamer Aldwairi

arXiv 2609.22284首次发表:更新:

发表机构

Temple University(天普大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究利用rPPG波形和唇部DCT系数两种轻量级视觉线索检测说话人脸深度伪造,通过融合达到AUC 0.891,并分析跨生成器迁移的互补性。

AI 中文摘要

基于rPPG的方法在不同生成器上的说话人脸(TF)深度伪造检测表现不均匀。我们研究了两种轻量级视觉线索:由RhythmFormer提取的rPPG导出波形和唇部区域离散余弦变换(DCT)系数,在Celeb-DF++的七种TF方法上,采用受试者独立协议进行评估。在域内,唇部区域DCT在除SadTalker外的所有方法上匹配或超过rPPG导出的1D ResNet,并且Concat融合达到AUC 0.891,而单模态基线分别为0.824和0.827。在留一生成器外评估下,这些线索出现分化:每种线索对三个保留方法迁移效果明显更好,而IP-LAP对两者均接近随机水平。Concat平均AUC为0.798,但在DCT迁移效果差的情况下低于仅使用rPPG,因此静态融合仅部分利用了这种互补性。唇部区域DCT在七种方法中的六种上优于全脸DCT。我们将rPPG导出的信号视为经验线索,并不声称其源于心脏活动。

英文摘要

Talking-face (TF) deepfakes are detected unevenly by rPPG-based methods across generators. We study two lightweight visual-only cues, rPPG-derived waveforms extracted by RhythmFormer and lip-region discrete cosine transform (DCT) coefficients, on the seven TF methods of Celeb-DF++ under a subject-independent protocol. In-domain, lip-region DCT matches or exceeds the rPPG-derived 1D ResNet on every method except SadTalker, and Concat fusion reaches AUC 0.891 against 0.824 and 0.827 for the unimodal baselines. Under leave-one-generator-out evaluation the cues split: each transfers clearly better to three held-out methods, and IP-LAP is near chance for both. Concat averages 0.798 but falls below rPPG alone where DCT transfers poorly, so static fusion only partly exploits this complementarity. Lip-region DCT outperforms full-face DCT on six of seven methods. We treat the rPPG-derived signal as an empirical cue and do not claim it is cardiac in origin.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑