arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

保留完全视觉保真度的音频驱动对抗防御3D说话人脸生成

Audio-Driven Adversarial Defense for 3D Talking Face Generation with totally Visual Fidelity Preservation

Rui-Qing Sun, Chen-Hao Cui, Hui-Yang Zhao, Tian Lan, Zhijing Wu, Xian-Ling Mao

arXiv 2608.30951首次发表:更新:

发表机构

Beijing Institute of Technology(北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对音频驱动3D说话人脸生成的隐私风险,提出一种基于心理声学掩蔽的不可感知音频扰动防御方法,可有效降低该生成效果并保持良好感知质量。

AI 中文摘要

生成式肖像模型的快速发展引发了人们对隐私泄露和身份滥用的日益担忧。特别是,音频驱动的3D说话人脸生成可以从单目视频中重建目标人物的可重复使用的3D肖像,并使用任意语音对其进行动画处理,使得逼真的身份模仿变得惊人地实用。现有的主动防御主要在视觉领域操作,通过在面部区域注入微小扰动来破坏身份获取。然而,这种扰动通常会损害视觉质量,因为人脸具有强烈的结构先验和社会敏感性,并且很容易被常见的现实世界变换(如调整大小)削弱。为了克服这些限制,我们提出了一种针对音频驱动3D说话人脸生成的不可感知音频防御,将保护从视觉模态转移到音频模态。具体来说,我们利用心理声学掩蔽将保护性扰动隐藏在语音信号的感知掩蔽频率区域内,从而减少感知失真,同时抑制可靠的面部动画。大量实验表明,所提出的方法有效降低了3D说话人脸生成的效果,同时保持了良好的感知质量。这些发现凸显了心理声学引导的音频扰动作为隐私保护肖像保护的实用且有前景的方向。

英文摘要

The rapid development of generative portrait models has raised growing concerns about privacy leakage and identity misuse. In particular, audio-driven 3D talking face generation can reconstruct a reusable 3D portrait of a target person from a monocular video and animate it with arbitrary speech, making realistic identity impersonation alarmingly practical. Existing proactive defenses mainly operate in the visual domain by injecting subtle perturbations into acial regions to disrupt identity acquisition. However, such perturbations often compromise visual quality due to the strong structural priors and social sensitivity of human faces, and are easily weakened by common real-world transformations such as resizing. To overcome these limitations, we propose an imperceptible audio defense for audio-driven 3D talking face generation by shifting protection from the visual modality to the audio modality. Specifically,we exploit psychoacoustic masking to hide protective perturbations within perceptually masked frequency regions of the speech signal, thereby reducing perceptual distortion while suppressing reliable facial animation. Extensive experiments demonstrate that the proposed method effectively degrades 3D talking face generation while preserving favorable perceptual quality. These findings highlight psychoacoustically guided audio perturbations as a practical and promising direction for privacy-preserving portrait protection.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑