发表机构
LIA, Avignon University; EDL(阿维尼翁大学LIA; EDL)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出多视角离散词元增强策略,利用多个SSL编码器生成互补训练视角,提升基于LLM的语音识别性能,在LibriSpeech上显著降低词错误率。
AI 中文摘要
离散语音词元为语音编码器与大语言模型之间提供了紧凑的接口,用于自动语音识别,但单一词元化系统对所选编码器仍然敏感。我们提出多视角离散词元增强,这是一种简单的策略,通过从固定的SSL编码器(如HuBERT、WavLM和MMS-300M)生成替代词元序列来增强每个训练话语。这些词元化被视为共享LLM解码器的互补训练视角,使其暴露于更多样化的离散语音表示,而无需多编码器推理。在测试时,模型可以使用单一编码器运行。在LibriSpeech上,该方法一致地改善了所有编码器相对于独立训练基线的性能,其中WavLM在test-clean上达到3.30%的词错误率(WER),在test-other上达到8.13%。预算匹配的对照实验表明,性能提升来自编码器多样性而非数据量。对多视角假设进行ROVER融合进一步将WER改善至3.03%和7.38%。
英文摘要
Discrete speech tokens provide a compact interface between speech encoders and large language models for automatic speech recognition, but single-tokenization systems remain sensitive to the chosen encoder. We propose multi-view discrete token augmentation, a simple strategy that augments each training utterance by generating alternative token sequences from fixed SSL encoders, such as HuBERT, WavLM, and MMS-300M. These tokenizations are treated as complementary training views for a shared LLM decoder, exposing it to more diverse discrete speech representations without requiring multi-encoder inference. At test time, the model can operate with a single encoder. On LibriSpeech, the approach consistently improves all encoders over independently trained baselines, with WavLM reaching 3.30% WER on test-clean and 8.13% on test-other. Budget-matched controls show that the gains come from encoder diversity rather than data volume. ROVER over multi-view hypotheses further improves WER to 3.03% and 7.38%.
Journal refIEEE SLT 2026