arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37104cs.CLcs.AI

后训练在多语言推理中改变了什么?

What Does Post-Training Change in Multilingual Reasoning?

  • University of Luxembourg(卢森堡大学)
  • Seafill Open-Source Community(Seafill 开源社区)

机构由 AI 辅助整理,请以论文原文为准。

Hongyang Li, Xiao Li, Caesar Wu, Grégoire Danoy, Pascal Bouvry

AI总结:

本研究审计Qwen3模型在多语言竞赛数学任务中的推理表现,发现语言是访问障碍,并揭示后训练各阶段瓶颈转移,提出从英语枢轴到可靠多语言推理的建设性路径。

AI中文摘要:

开源推理模型在不同语言间提供不均衡的推理能力。当模型能够解决问题却无法以用户的语言提供完整解答时,语言便成为访问障碍,而不仅仅是性能差异的来源。我们对Qwen3检查点在十一种语言的竞赛数学任务上进行了审计。在十种非英语语言中,仅有15.4%-17.9%的问题在16次采样中的任意一次获得了以所请求语言呈现的、带有可见推理的正确且终止的解答,而英语中这一比例为92.9%。为确定这一差异的来源,我们评估了来自同一模型家族的十三个端点,涵盖已发布检查点、两种规模的多语言监督微调(SFT)、受控SFT消融实验以及三种强化学习(RL)奖励公式。我们联合跟踪正确性、语言遵循、终止性和交付效率。主要瓶颈在后训练各阶段间发生转移。已发布模型常以英语进行推理。多语言SFT恢复了目标语言推理,但多语言、仅英语和单语言SFT运行中的准确率均下降,表明这一代价并非多语言混合所特有;非英语推理轨迹还易陷入不终止的循环。RL在两种分支中均以零准确率代价恢复了终止性,但只有奖励包含语言项的分支实现了交付:仅奖励正确性会使模型回归英语。综合来看,这些阶段构建了一条从英语枢轴能力到可靠交付的多语言推理的建设性后训练路径。

英文摘要:

Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this disparity, we evaluate thirteen endpoints from one model family, spanning released checkpoints, multilingual supervised fine-tuning (SFT) at two scales, controlled SFT ablations, and three reinforcement-learning (RL) reward formulations. We jointly track correctness, language adherence, termination, and delivery efficiency. The dominant bottleneck shifts across post-training stages. Released models often reason in English. Multilingual SFT restores target-language reasoning, but accuracy declines across multilingual, English-only, and single-language SFT runs, showing that this cost is not specific to multilingual mixing; non-English reasoning traces additionally become prone to non-terminating loops. RL restores termination in both arms at no cost in accuracy, but only the arm whose reward includes a language term delivers: rewarding correctness alone returns the model to English. Together, these stages establish a constructive post-training path from English-pivoted capability to multilingual reasoning that is reliably delivered.

补充信息

↑