arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越编码器融合:基于多视角离散词元增强的LLM语音识别

Beyond Encoder Fusion: Multi-View Discrete Token Augmentation for LLM-Based ASR

Paul Moïse Gangbadja, Mickael Rouvier, Fabrice Lefèvre

arXiv 2609.23525首次发表:更新:

发表机构

LIA, Avignon University; EDL(阿维尼翁大学LIA; EDL)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出多视角离散词元增强策略,利用多个SSL编码器生成互补训练视角,提升基于LLM的语音识别性能,在LibriSpeech上显著降低词错误率。

AI 中文摘要

离散语音词元为语音编码器与大语言模型之间提供了紧凑的接口,用于自动语音识别,但单一词元化系统对所选编码器仍然敏感。我们提出多视角离散词元增强,这是一种简单的策略,通过从固定的SSL编码器(如HuBERT、WavLM和MMS-300M)生成替代词元序列来增强每个训练话语。这些词元化被视为共享LLM解码器的互补训练视角,使其暴露于更多样化的离散语音表示,而无需多编码器推理。在测试时,模型可以使用单一编码器运行。在LibriSpeech上,该方法一致地改善了所有编码器相对于独立训练基线的性能,其中WavLM在test-clean上达到3.30%的词错误率(WER),在test-other上达到8.13%。预算匹配的对照实验表明,性能提升来自编码器多样性而非数据量。对多视角假设进行ROVER融合进一步将WER改善至3.03%和7.38%。

英文摘要

Discrete speech tokens provide a compact interface between speech encoders and large language models for automatic speech recognition, but single-tokenization systems remain sensitive to the chosen encoder. We propose multi-view discrete token augmentation, a simple strategy that augments each training utterance by generating alternative token sequences from fixed SSL encoders, such as HuBERT, WavLM, and MMS-300M. These tokenizations are treated as complementary training views for a shared LLM decoder, exposing it to more diverse discrete speech representations without requiring multi-encoder inference. At test time, the model can operate with a single encoder. On LibriSpeech, the approach consistently improves all encoders over independently trained baselines, with WavLM reaching 3.30% WER on test-clean and 8.13% on test-other. Budget-matched controls show that the gains come from encoder diversity rather than data volume. ROVER over multi-view hypotheses further improves WER to 3.03% and 7.38%.

Journal refIEEE SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑