Ariadne的唇形同步之线:通过唇部运动与头部姿态之间的不一致性揭示伪造视频
Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses
- University of Science and Technology of China(中国科学技术大学)
- Shanghai Jiaotong University(上海交通大学)
- Peking University(北京大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对唇形同步伪造视频,提出LipDA框架,利用唇部运动与头部姿态的不一致性进行检测与归因,在检测和模型归因上分别取得超97% AUC和97.5%准确率。
AI中文摘要:
唇形同步生成技术的最新进展已能创造出高度逼真的视频,对社会构成严重风险。然而,现有的防御策略难以应对唇形同步伪造,因为先进的唇形同步生成方法不仅实现了更好的唇部同步,还消除了视觉伪影。一个重要原因是它们忽视了自然语音视频中唇部运动与头部姿态之间固有的生物耦合。本文提出LipDA,一个用于联合唇形同步检测与归因的新框架,利用头部与唇部之间的不一致性。对于检测,该框架通过对比真实与伪造视频中的唇部和姿态特征来学习量化这种差异。对于归因,我们的方法旨在捕捉作为模型指纹的独特时间动态和音视频同步模式,从而实现来源追踪。我们在两个具有挑战性的唇形同步数据集以及我们自行提出的大规模多生成器数据集上进行了广泛实验。LipDA在检测上实现了超过97%的AUC,在模型归因上达到97.5%的准确率,显著优于现有方法。代码和所提出的LipSync-A数据集可在以下https URL获取。
英文摘要:
Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as advanced LipSync generation methods not only achieve better lip synchronization but also eliminate visual artifacts. An important reason is that they overlook an inherent biological coupling between lip movements and head poses in natural speech videos. In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip. For detection, the framework learns to quantify this discrepancy by contrasting lip and pose features from authentic versus forged videos. For attribution, our method is designed to capture the unique temporal dynamics and audio-visual synchronization patterns that act as the fingerprint of models, enabling source tracing. We conduct extensive experiments on two challenging LipSync datasets as well as our own proposed large-scale and multi-generator dataset. LipDA achieves over 97\% AUC in detection and 97.5\% accuracy in model attribution, significantly outperforming existing methods. Code and the proposed LipSync-A dataset are available at https://github.com/AnsonShe/LipDA.