发表机构
Kyoto University(京都大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出通过模型合并将外部语言模型直接集成到基于LLM的ASR参数中,避免推理额外成本,在CSJ和LibriSpeech上验证了目标领域性能提升且不损速度与内存。
AI 中文摘要
自动语音识别(ASR)系统在配对语音-文本数据上训练,通过利用仅在文本数据上训练的语言模型(LMs)得到了改进。浅层融合和密度比等LM融合方法是在ASR解码过程中整合外部LM的成熟方法。然而,由于LM推理,它们产生了额外的计算成本,这对于近期更大的LM尤其成问题。在本研究中,我们提出通过模型合并来整合外部LM。该方法将LM直接集成到基于LLM的ASR模型的参数中,在推理时无需额外计算成本。我们通过对LoRA参数进行算术运算来公式化领域扩展和迁移。实验评估针对在CSJ和LibriSpeech上训练的基于LLM的ASR的领域适应进行。我们表明,我们的LM合并持续提高了目标领域的ASR性能,且不降低推理速度或内存占用。
英文摘要
Automatic speech recognition (ASR) systems, trained on paired speech-text data, have been improved by leveraging language models (LMs) trained on text-only data. LM fusion methods such as shallow fusion and density ratio are well-established methods that incorporate external LMs during ASR decoding. However, they incur additional computational costs due to LM inference, which is particularly problematic for recent larger LMs. In this study, we propose incorporating external LMs via model merging. This method integrates the LMs directly into the parameters of an LLM-based ASR model, requiring no additional computational cost at inference. We formulate domain extension and transfer via arithmetic operations on LoRA parameters. Experimental evaluations were conducted for the domain adaptation of LLM-based ASR trained on CSJ and LibriSpeech. We show that our LM merging consistently improved the ASR performance in the target domains, without degrading inference speed or memory footprint.
CommentsAccepted to Interspeech2026