arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越短片段:利用向量档案扩展说话人嵌入

Beyond Short Segments : Expanding Speaker Embeddings with Vector Archives

Hyunku Kang, Minkyu Cho, Chanwoo Kim

arXiv 2609.25007首次发表:更新:

发表机构

Korea University(高丽大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对短语音说话人验证性能下降问题,提出VAM-ECAPA系统,利用基于Transformer的TVAMSP模块将稀疏特征映射到可学习向量档案,在VoxCeleb1上1秒片段取得8.334% EER,相对错误率降低54.8%。

AI 中文摘要

最先进的说话人验证(SV)系统在短语音上的性能会因说话人特定信息不足而严重下降。为应对这一关键挑战,我们提出了向量档案映射ECAPA(VAM-ECAPA),一种旨在增强短时语音特征提取的新型系统。我们系统的核心是基于Transformer的统计池化向量档案映射(TVAMSP)模块,该模块通过将信息稀缺的特征与可学习的典型说话人特征向量档案进行映射来丰富这些特征。通过将TVAMSP模块集成到强大的WavLM+ECAPA-TDNN基线中,我们的系统学会将短片段中的稀疏特征映射为稳健且具有判别性的说话人表示。在VoxCeleb1基准上的实验表明,我们提出的VAM-ECAPA在1秒测试片段上实现了极具竞争力的8.334%等错误率(EER),与传统训练的基线相比,相对错误率降低了54.8%。

英文摘要

The performance of state-of-the-art speaker verification (SV) systems severely degrades on short utterances due to insufficient speaker-specific information. To address this critical challenge, we propose the Vector Archive Mapping ECAPA (VAM-ECAPA), a novel system designed to enhance feature extraction from short-duration speech. The core of our system is the Transformer-based Vector Archive Mapping with Statistical Pooling (TVAMSP) module, which enriches information-scarce features by mapping them against a learnable Vector Archive of canonical speaker traits. By integrating the TVAMSP module into a strong WavLM+ECAPA-TDNN baseline, our system learns to map sparse features from short segments into robust, discriminative speaker representations. Experiments on the VoxCeleb1 benchmark show that our proposed VAM-ECAPA achieves a highly competitive EER of 8.334% on 1-second test segments, a 54.8% relative error reduction compared to a conventionally-trained baseline.

CommentsAccepted at INTERSPEECH 2026 (oral)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑