MoLGE:用于大规模多语言语音识别高效扩展的语言组专家混合模型
MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition
浏览论文内容
中文总结 AI 辅助
研究多语言ASR模型面临的多语言诅咒问题,提出MoLGE模型,基于S3Ms为相似语言簇分配专家模块并集成LoRA策略,研究语言分组策略影响,实验表明该模型以最少参数提升性能,为扩展语言覆盖提供有效途径。
中文摘要 AI 辅助
涵盖数百种语言的大规模多语言自动语音识别(ASR)模型必须在不同语言和声学条件下保持稳健性能。然而,这些模型常面临多语言诅咒,模型能力在各语言间被稀释。为应对这一挑战,我们提出基于语音自监督模型(S3Ms)构建的语言组专家混合模型(MoLGE)。MoLGE为相似语言簇分配专用专家模块,减少所需子模块数量。它还将分层低秩适应(LoRA)策略集成到S3M架构的解耦声学和语言组件中,在保持参数效率的同时实现特定语言特征的高效建模。此外,我们研究了基于语言和数据驱动标准的语言分组策略对整体性能的影响。在实验中,我们在包含495种语言的多语言基准上评估MoLGE。结果表明,MoLGE以最少的可训练参数增加持续优于密集多语言基线。这些语言分组策略在ASR建模的语音和拼写方面都有显著改进。我们的发现表明,结构化语言专业化为大规模扩展多语言ASR的语言覆盖范围提供了有效途径。
英文摘要
Massively multilingual automatic speech recognition (ASR) models covering hundreds of languages must maintain robust performance across diverse linguistic and acoustic conditions. However, these models often encounter the curse of multilinguality, where model capacity is diluted across languages. To address this challenge, we propose Mixture of Language Group Experts (MoLGE), built upon speech self-supervised models (S3Ms). MoLGE assigns dedicated expert modules to clusters of similar languages, reducing the number of required submodules compared to conventional language-specific Mixture-of-Experts (MoE) schemes. It further integrates a hierarchical Low-Rank Adaptation (LoRA) strategy into the disentangled acoustic and linguistic components of the S3M architecture, enabling efficient modeling of language-specific characteristics while maintaining parameter efficiency. Further, we investigate the impact of language grouping strategies based on both linguistic and data-driven criteria on overall performance, providing an interpretable perspective on how language structure influences scalability in multilingual speech systems. In experiments, we evaluate MoLGE on a multilingual benchmark encompassing 495 languages. Results demonstrate that MoLGE consistently outperforms dense multilingual baselines with a minimal increase in trainable parameters. Notably, these language grouping strategies yield substantial improvements for both phonetic and orthographic aspects of ASR modeling. Our findings suggest that structured language specialization provides an effective pathway for massively scaling language coverage of multilingual ASR.
发表机构
- Dept. of Electronics and Electrical Engineering, Yonsei University(延世大学电子与电气工程系)
机构由 AI 辅助整理,请以论文原文为准。