arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23650cs.SDeess.AS

揭露你的伪装:从语音转换中恢复源说话者身份

Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion

Hanlei Zhang, Zhongming Ma, Mingyang Zhang, Tengfei Liu, Yushi Cheng, Yanjiao Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对语音转换威胁生物特征安全问题,提出TRIDENT框架,通过三叉戟架构,包括识别转换机制、提取目标说话者潜在表示等,从转换音频恢复源说话者身份,实验表明其准确率高且性能稳健。

中文摘要 AI 辅助

语音转换(VC)通过允许攻击者冒充目标说话者,对生物特征安全构成重大威胁。在法医环境中,从转换后的音频中恢复源说话者的身份对于缩小嫌疑人范围至关重要。为了解决这个问题,我们提出了TRIDENT,一个旨在从转换后的音频样本中恢复源说话者原始身份的追溯框架。TRIDENT采用了一个由主提取器和两个辅助分支组成的三叉戟架构。第一个辅助分支识别潜在的语音转换机制。这种设计承认,即使确切的转换策略未知,攻击者采用的高性能模型通常是既定主流模型的衍生或变体。第二个辅助分支提取目标说话者的潜在表示,便于从复合转换后的音频样本中分离出目标特定特征。最后,主提取器利用两个辅助分支的见解来解耦混杂因素,并提炼出源说话者身份的高度判别性表示。实验结果表明,TRIDENT在对抗7种先进的语音转换方法时,准确率高达90.99%。此外,TRIDENT在具有挑战性的条件下保持稳健的性能,包括电话信道、未见语言和自适应场景。

英文摘要

Voice conversion (VC) poses a significant threat to biometric security by allowing attackers to impersonate target speakers. In forensic contexts, recovering the source speaker's identity from converted audio is vital for narrowing the field of suspects. To address this, we propose TRIDENT, a retracing framework designed to restore a source speaker's original identity from a converted audio sample. TRIDENT utilizes a three-pronged architecture consisting of a primary extractor and two auxiliary branches. The first auxiliary branch identifies the underlying voice conversion mechanism. This design acknowledges that even if the exact conversion strategy is unknown, a high-performance model adopted by the attacker is typically a derivative or variant of established mainstream ones. The second auxiliary branch extracts a latent representation of the target speaker, facilitating the isolation of target-specific traits from the composite converted audio sample. Finally, the main extractor leverages insights from both auxiliary branches to decouple confounding factors and distill a highly discriminative representation of the source speaker's identity. Experimental results demonstrate that TRIDENT achieves an accuracy as high as 90.99% against 7 state-of-the-art voice conversion methods. Furthermore, TRIDENT maintains robust performance under challenging conditions, including telephony channels, unseen languages, and adaptive scenarios.

发表机构

  • Zhejiang University(浙江大学)
  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

↑