发表机构
Nankai University; Xingchen AGI Lab, China Telecom Artificial Intelligence Technology Co. Ltd(南开大学; 星辰AGI实验室,中国电信人工智能技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语音语言模型中ASR专业化损害QA能力的问题,提出任务特定在线策略蒸馏,利用ASR前后模型分别监督任务轨迹,实验证明兼顾识别与QA性能且稳健。
AI 中文摘要
语音语言模型(SLMs)从预训练语言模型中继承了强大的指令遵循能力,然而ASR专业化可能会大幅削弱这些能力。为了解决这种ASR与QA之间的权衡问题,我们提出了任务特定在线策略蒸馏(TS-OPD),该方法利用ASR专业化前后的模型作为互补的QA和ASR教师。学生模型为ASR和QA生成各自的任务条件轨迹,每个轨迹仅由对应的教师监督,从而减少两种监督信号之间的直接竞争。在基础ASR、上下文ASR和QA上的实验表明,TS-OPD在提升识别性能的同时保持了QA能力。此外,TS-OPD在不同平衡系数下保持稳健,并继续从增加的蒸馏数据中获益。
英文摘要
Speech Language Models (SLMs) inherit strong instruction-following capabilities from pretrained language models, yet ASR specialization can substantially degrade them. To address this ASR--QA trade-off, we propose Task-Specific On-Policy Distillation (TS-OPD), which leverages models before and after ASR specialization as complementary QA and ASR teachers. The student generates separate task-conditioned trajectories for ASR and QA, each supervised only by its corresponding teacher, thereby reducing direct competition between the two supervision signals. Experiments on basic ASR, contextual ASR, and QA demonstrate that TS-OPD improves recognition while preserving QA capability. Moreover, TS-OPD remains robust across different balancing coefficients and continues to benefit from increased distillation data.