语音令牌会泄露声纹吗?针对端到端语音语言模型的说话人反转攻击
Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
浏览论文内容
中文总结 AI 辅助
研究端到端语音语言模型中语音令牌是否泄露声纹,提出用Audio BERT构建令牌嵌入及SpInv两阶段反转方法,在VoxCeleb数据集实验显示,仅三秒前端输出,SpInv在指定空间能达高于0.70的余弦相似度。
中文摘要 AI 辅助
端到端语音语言模型越来越多地用语音令牌来表示用户语音,而非仅依赖级联的ASR-LLM-TTS管道。尽管这些令牌支持富有表现力且低延迟的口语交互,但可能保留敏感的说话人特征。我们研究暴露的语音令牌是否会泄露声纹,并将此风险表述为说话人反转攻击。我们引入了Audio BERT(AuB),一种可训练模型,从离散码本构建令牌嵌入并将其聚合为对说话人敏感的表示,还提出了SpInv,一种基于AuB的两阶段反转方法,用于在攻击者指定的说话人编码器空间中恢复嵌入。我们在VoxCeleb数据集上使用说话人不相交协议评估了Moshi、Higgs3、Kimi-Audio和Qwen3-Omni。大量实验表明,仅使用三秒的前端输出,SpInv在攻击者指定的说话人编码器空间中就能实现高于0.70的余弦相似度。
英文摘要
End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens support expressive and low-latency spoken interaction, they may also preserve sensitive speaker characteristics. We investigate whether exposed speech tokens leak voiceprints and formulate this risk as a speaker inversion attack. We introduce Audio BERT (AuB), a trainable model that constructs token embeddings from discrete codebooks and aggregates them into speaker-sensitive representations, and propose SpInv, a two-stage inversion method built on AuB to recover embeddings in the space of an attacker-specified speaker encoder. We evaluate Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni using speaker-disjoint protocols on the VoxCeleb dataset. Extensive experiments show that, with only three seconds of frontend output, SpInv achieves cosine similarities above 0.70 in the attacker-specified speaker-encoder space.