arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效的多语言推理迁移通过渐进式代码切换

Efficient Multilingual Reasoning Transfer via Progressive Code-Switching

Zhijun Wang, Junxiao Liu, Hao Zhou, Hao-Ran Wei, Baosong Yang, Shujian Huang

arXiv 2607.00485首次发表:更新:

发表机构

Tongyi Lab(通义实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出渐进式代码切换(PCS)框架,通过轻量翻译和逐步强化学习,将英语推理能力高效迁移到目标语言,缩小性能差距。

AI 中文摘要

大型推理模型(LRMs)在英语中表现出强大的推理能力,但在需要用其他语言推理时性能显著下降。一个自然的解决方案是将模型的英语推理能力迁移到目标语言。然而,现有的迁移方法通常依赖于从更强的LRMs中提取的目标语言推理轨迹或外部评判模型的在线监督,这些方法成本高昂且难以扩展。在本文中,我们提出PCS(渐进式代码切换),一种更高效的迁移框架,仅需要轻量级翻译,无需任何更强的模型进行蒸馏或评判。PCS首先通过将一部分英语推理步骤翻译成目标语言来构建代码切换的推理轨迹,并通过监督微调初始化模型的代码切换能力。然后,它应用带有步骤级语言一致性课程表的强化学习,逐步提高目标语言比例,直到模型完全用目标语言推理。这种渐进式设计提供了平滑的迁移路径,避免了直接强制目标语言推理时常见的性能不稳定和下降。在多个基准测试和五种类型多样的语言上的实验表明,PCS显著缩小了目标语言与英语推理之间的性能差距,在保持竞争性准确性的同时产生更一致的语言推理。

英文摘要

Large reasoning models (LRMs) have achieved strong reasoning capabilities in English, yet their performance degrades significantly when required to reason in other languages. A natural solution is to transfer the model's English reasoning ability to target languages. However, existing transfer approaches typically rely on distilled target-language reasoning traces from stronger LRMs or online supervision from external judge models, which are costly and difficult to scale. In this paper, we propose PCS (Progressive Code-Switching), a more efficient transfer framework that requires only lightweight translation without any stronger model for distillation or judging. PCS first constructs code-switched reasoning traces by translating a subset of English reasoning steps into the target language, and uses them to initialize the model's code-switching ability via supervised fine-tuning. It then applies reinforcement learning with a step-level language consistency curriculum, progressively raising the target-language ratio until the model reasons entirely in the target language. This progressive design provides a smooth transfer path that avoids the instability and performance degradation commonly observed when directly enforcing target-language reasoning. Experiments on multiple benchmarks and five typologically diverse languages show that PCS substantially narrows the performance gap between target-language and English reasoning, yielding more language-consistent reasoning while maintaining competitive accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑