arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生理信号作为一种用于检测说话人脸深度伪造的法医模态

Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection

Othmane Harraq, Tamer Aldwairi

arXiv 2607.21776首次发表:更新:

发表机构

Temple University(天普大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对说话人脸深度伪造检测难题,提出用RhythmFormer提取rPPG波形并训练轻量级分类器的框架。在Celeb-DF++的TF子集上评估,1D ResNet表现良好,还揭示检测难度因方法而异,体现生成器生理特性,这是主要理论贡献。

AI 中文摘要

说话人脸(TF)深度伪造生成技术能从静态源图像和音频信号合成逼真的面部视频,现有基于图像的检测器难以识别。与面部交换不同,TF合成没有真实视频可继承生理特征,使得远程光电容积脉搏波描记法(rPPG)成为检测此类伪造的独特方法。我们提出一个检测框架,通过RhythmFormer提取每个视频的rPPG波形,并训练一组轻量级分类器区分真实和合成的生理信号。在Celeb-DF++的TF子集上严格独立于主体的协议下评估,我们的1D ResNet实现了0.806的AUC和27.8%的EER,与最佳通用检测器相差2.4个点。我们记录了对代表性的先前rPPG检测器DeepFakesON-Phys的控制再现研究,其在传统面部交换数据上的AUC为0.999,在Celeb-DF++的TF子集上降至0.622。我们还表明检测难度强烈依赖方法:七个TF生成器的AUC范围从0.985(Real3DPortrait)到0.690(IP-LAP),排名在所有评估协议中保持稳定。这种差异反映了每个生成器可解释的生理特性,而非评估噪声,是该工作的主要理论贡献。

英文摘要

Talking-face (TF) deepfake generation synthesizes photore- alistic facial video from a static source image and an au- dio signal, producing forgeries that current image-based detectors consistently fail to identify. Unlike face-swap ma- nipulation, TF synthesis has no underlying real video from which to inherit physiological characteristics, making re- mote photoplethysmography (rPPG) a uniquely motivated detection modality for this forgery category. We propose a detection framework that extracts per-video rPPG wave- forms via RhythmFormer and trains a suite of lightweight classifiers to distinguish real from synthesized physiologi- cal signals. Evaluated on the TF subset of Celeb-DF++ un- der a strict subject-independent protocol, where test identi- ties are completely separated from training identities, our 1D ResNet achieves an AUC of 0.806 and EER of 27.8%, placing it within 2.4 points of the best published general- purpose detector (Effort, ICML 2025) while operating ex- clusively on the physiological channel. We document a con- trolled reproduction study of DeepFakesON-Phys, the rep- resentative prior rPPG detector, demonstrating degrada- tion from AUC 0.999 on legacy face-swap data to 0.622 on the TF subset of Celeb-DF++. We further show that detec- tion difficulty is strongly method-dependent: AUC ranges from 0.985 (Real3DPortrait) to 0.690 (IP-LAP) across the seven TF generators, with the ranking remaining perfectly stable across all evaluation protocols. This spread reflects an interpretable physiological property of each generator rather than evaluation noise, and constitutes the primary theoretical contribution of the work.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑