arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用离散音频令牌的文本无关说话人验证

Text-Independent Speaker Verification Using Discrete Audio Tokens

Zheng Liang, Junjie Li, Kong Aik Lee

arXiv 2607.07579首次发表:更新:

AI 中文总结

研究在自动说话人验证中神经音频编解码器离散表示性能不佳的问题,提出跨特征知识蒸馏框架,通过引导学生模型模仿教师模型嵌入空间,有效利用令牌中说话人信息,提升了基于编解码器系统的ASV性能。

AI 中文摘要

神经音频编解码器(NAC)能实现高效音频压缩并在语音合成等下游任务取得成功,但在自动说话人验证(ASV)中,其离散表示性能逊于传统频谱特征。经实证表明,说话人线索隐含于离散令牌中却未被传统ASV训练范式充分利用。为此提出跨特征知识蒸馏(CFKD)框架,引导基于编解码器的学生模型模仿强大的基于Fbank的教师模型的嵌入空间,为有效利用令牌中的说话人信息提供结构化监督。VoxCeleb基准测试表明,CFKD显著提升基于编解码器系统的ASV性能,使其接近基于Fbank的教师模型的准确率,凸显离散音频令牌在多样语音任务中的潜力。

英文摘要

Neural audio codecs (NACs) enable efficient audio compression and have achieved success in downstream tasks such as speech synthesis. However, their discrete representations consistently underperform traditional spectral features in automatic speaker verification (ASV). We empirically demonstrate that speaker cues are implicitly preserved in discrete tokens but remain underutilized by conventional ASV training paradigms. To address this, we propose a Cross-Feature Knowledge Distillation (CFKD) framework. By guiding the codec-based student to mimic the embedding space of a strong Fbank-based teacher, CFKD provides structured supervision for effective utilization of speaker information in tokens. Experiments on the VoxCeleb benchmarks show that CFKD substantially improves the ASV performance of codec-based systems, allowing them to approach the accuracy of Fbank-based teacher models and highlighting the potential of discrete audio tokens for diverse speech tasks.

CommentsThis paper has been accepted by Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑