面向多语言数学推理的在线策略增量蒸馏
On-Policy Delta Distillation for Multilingual Math Reasoning
- NAVER AI Lab(NAVER AI实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对多语言数学推理任务,提出OPD²方法,实验表明其在韩语、日语上性能提升显著,且能缩小英语与韩语的性能差距,凸显多语言数据的重要性。
AI中文摘要:
在线策略蒸馏(OPD)正成为大语言模型(LLM)后训练中替代强化学习的有前景方案,但其在多语言场景的有效性仍未得到充分探索。本研究针对英语、韩语、日语的数学推理任务,探究OPD及其高级变体在线策略增量蒸馏(OPD²)。OPD²通过利用后训练教师模型与基础模型间的概率差距作为学习信号,对OPD进行改进。基于Qwen3的实验表明,OPD²始终优于原始OPD,在韩语和日语上提升尤为显著,还总体缩小了英语与韩语间的性能差距。研究进一步发现,仅针对英语的OPD也能提升韩语和日语的性能,但常使模型响应偏向英语,凸显了多语言数据对保留目标语言响应的重要性。
英文摘要:
On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. OPD$^2$ improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD$^2$ consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.