面向多方言自动语音识别的策略内自蒸馏:掌握方言,保留普通话
On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin
AI总结:
本研究针对ASR模型方言识别效果有限且直接适配易降低普通话准确率的问题,提出策略内自蒸馏(OPSD)方法,结合CPT、SFT优化Qwen3-ASR-1.7B,在不降低普通话识别能力的同时提升了多方言识别效果。
AI中文摘要:
近期的大规模自动语音识别(ASR)模型已实现出色的普通话识别准确率,且具备一定的中文方言识别能力,但在实际场景语音中,其方言识别准确率仍有限。直接的方言适配可降低方言的字符错误率(CER),却可能提升普通话的CER。因此,本研究探讨如何适配高性能ASR模型,在不降低普通话识别能力的前提下提升多方言识别效果。我们采用的适配流程为:持续预训练(CPT)与方言监督微调(SFT)提供坚实基础,策略内自蒸馏(OPSD)作为最终优化模块。OPSD通过让学生模型在自身解码的前缀上训练,同时以冻结的教师模型(以参考 transcript 作为特权上下文)提供软 token 级目标,解决自回归ASR中的训练-测试不匹配问题。该方法用蒸馏替代方言数据上的硬交叉熵更新,在优化方言识别的同时保留普通话能力。我们以Qwen3-ASR-1.7B实例化该框架,并在公开及内部的普通话与方言测试集上评估。在匹配的优化数据与 schedule 下,OPSD提升了方言识别准确率且未提高普通话CER,而持续的教师强制微调则会提升普通话CER。我们将发布模型权重与评估脚本。
英文摘要:
Recent large-scale ASR models already achieve strong Mandarin recognition accuracy and have some ability to recognize Chinese dialects. However, their dialect recognition accuracy is still limited in real-world speech. Direct dialect adaptation can lower dialect CER, but it may also raise Mandarin CER. We therefore study how to adapt a capable ASR model to improve multi-dialect recognition without degrading Mandarin recognition. We adopt an adaptation pipeline where continual pre-training (CPT) and dialect supervised fine-tuning (SFT) provide a strong foundation, and On-Policy Self-Distillation (OPSD) serves as the final refinement. OPSD addresses the train--test mismatch in autoregressive ASR by training the student model on its own decoded prefixes while a frozen teacher, conditioned on the reference transcript as privileged context, provides soft token-level targets. This replaces hard cross-entropy updates on dialect data with distillation, preserving Mandarin ability while refining dialect recognition. We instantiate the framework with Qwen3-ASR-1.7B and evaluate it on public and internal Mandarin and dialect test sets. Under matched refinement data and schedule, OPSD improves dialect recognition without raising Mandarin CER, whereas continued teacher-forced fine-tuning increases Mandarin CER. We will release the model weights and evaluation scripts.