arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考语音-LLM集成用于ASR:通过交错实现有效的联合语音-文本训练

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren, Keqi Deng, Xiaoyang Chen, Ali Zare, Bo Ren, Yuxuan Hu, Junkun Chen, Yan Huang, Yelong Shen, Jinyu Li

arXiv 2607.01733首次发表:更新:

发表机构

Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出JSTIP,一种面向ASR的预训练策略,通过构建词级和段级交错语音-文本序列,在38k小时数据上相比基线提升了实体识别准确率,并缩小了模态差距。

AI 中文摘要

语音-LLM集成通过利用大量文本预训练显示出有希望的结果,但其对自动语音识别(ASR)的具体益处仍不明确。我们观察到,随着监督ASR训练数据的增加,LLM先验的贡献变得不那么明显,而简单的语音-文本联合训练未能充分利用文本知识。因此,我们提出联合语音-文本交错预训练(JSTIP),这是一种面向ASR的预训练策略,针对接受连续输入的语音-LLM架构,在对齐对中构建词级和段级交错的语音-文本序列。在38k小时的ASR数据上的实验表明,与仅ASR和联合语音-文本训练基线相比,实体准确率持续提升。JSTIP使用领域转录文本即可达到与合成语音-文本对相当的实体识别性能,简化了领域适应。得益于文本预训练和领域文本数据,JSTIP在医学实体识别方面与开源ASR和语音-LLM系统具有竞争力。零样本语音问答行为进一步表明,交错减少了语音-文本模态差距并保留了LLM生成先验,这很可能是ASR任务上实体改进的原因。

英文摘要

Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data increases, the contribution of LLM priors becomes less evident, and simple speech-text joint training under-utilizes textual knowledge. We therefore propose Joint Speech-Text Interleaved Pretraining (JSTIP), an ASR-oriented pretraining strategy that constructs word-level and segment-level interleaved speech-text sequences within aligned pairs for speech-LLM architectures that accept continuous inputs. Experiments on 38k hours of ASR data show consistent entity accuracy improvement compared to ASR-only and joint speech-text training baselines. JSTIP achieves on-par entity recognition performance using domain transcription text compared to synthetic speech-text pairs, simplifying domain adaptation. Benefiting from textual pretraining and domain text data, JSTIP is competitive with open-source ASR and Speech-LLM systems in medical entity recognition. The zero-shot speech question answering behaviors further suggest that interleaving reduces the speech-text modality gap and preserves the LLM generative prior, which is likely the reason for the entity improvements on the ASR task.

CommentsAccepted to SLT 2026, camera-ready version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑