arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05802cs.CLcs.LG

面向多语言数学推理的在线策略增量蒸馏

On-Policy Delta Distillation for Multilingual Math Reasoning

Byeongho Heo, Jaehui Hwang, Sangdoo Yun, Dongyoon Han

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对多语言数学推理任务,提出OPD²方法,实验表明其在韩语、日语上性能提升显著,且能缩小英语与韩语的性能差距,凸显多语言数据的重要性。

中文摘要 AI 辅助

在线策略蒸馏(OPD)正成为大语言模型(LLM)后训练中替代强化学习的有前景方案,但其在多语言场景的有效性仍未得到充分探索。本研究针对英语、韩语、日语的数学推理任务,探究OPD及其高级变体在线策略增量蒸馏(OPD²)。OPD²通过利用后训练教师模型与基础模型间的概率差距作为学习信号,对OPD进行改进。基于Qwen3的实验表明,OPD²始终优于原始OPD,在韩语和日语上提升尤为显著,还总体缩小了英语与韩语间的性能差距。研究进一步发现,仅针对英语的OPD也能提升韩语和日语的性能,但常使模型响应偏向英语,凸显了多语言数据对保留目标语言响应的重要性。

英文摘要

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. OPD$^2$ improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD$^2$ consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.

发表机构

  • NAVER AI Lab(NAVER AI实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑