面向多语言大语言模型的ASR的语言专业化多教师在线策略蒸馏
Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR
中文总结 AI 辅助
针对多语言ASR的跨语言优化冲突问题,提出LS-MOPD方法,通过语言专业化教师与多语言学生的在线策略蒸馏,在多语言基准测试中实现了优于RL基线及最优教师的性能。
中文摘要 AI 辅助
现代基于大语言模型(LLM)的自动语音识别(ASR)系统已将多语言能力确立为标准特征,其利用大规模多语言语料库和LLM的跨语言知识,在多语言基准测试中实现了有竞争力的性能。然而,对声学、音系和词汇特征异质的语言进行联合建模,不可避免地会引入优化冲突,损害各语言的专业化程度。为解决这一挑战,本文提出了语言专业化多教师在线策略蒸馏(LS-MOPD)方法,该方法将语言特定知识的获取与多语言能力的解耦:语言专业化教师通过强化学习(RL)独立优化,之后通过语言路由和词级别多教师蒸馏将其专业知识整合到通用多语言学生模型中,从而减少直接跨语言优化冲突。本文进一步探索了两种声学前缀配置:静态和动态,以研究师生前缀一致性如何影响在线策略蒸馏的效果。在涵盖普通话、普通话方言、粤语和英语的基准测试上进行的实验表明,LS-MOPD的性能显著优于RL基线,且始终超过表现最佳的RL教师所定义的经验性能包络,揭示了其在多语言ASR中泛化到所有教师之外的潜力。
英文摘要
Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achieve competitive performance across multilingual benchmarks. However, jointly modeling languages with heterogeneous acoustic, phonological, and lexical characteristics inevitably introduces optimization conflicts, undermining language-wise specialization. To address this challenge, we propose Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD), which decouples language-specific knowledge acquisition from multilingual capability integration: language-specialized teachers are independently optimized via reinforcement learning (RL), with their expertise then integrated into a generalist multilingual student through language routing and token-level multi-teacher distillation, thereby reducing direct cross-lingual optimization conflicts. We further explore static and dynamic acoustic-prefix configurations to examine how teacher-student prefix consistency influences the efficacy of on-policy distillation. Experiments on benchmarks covering Mandarin, Mandarin subdialects, Cantonese, and English demonstrate that LS-MOPD substantially outperforms RL baselines and surpasses the empirical performance envelope defined by the best-performing RL teachers on nearly all benchmarks, revealing its potential to generalize beyond all teachers in multilingual ASR.