arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越英语的GRPO:非英语与多语言场景下GRPO的大规模研究

GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

Konstantin Dobler, Federico Scozzafava, Jonathan Janke, Mohamed Ali, Simon Lehnerer

arXiv 2608.13698首次发表:更新:

发表机构

Apple; Hasso Plattner Institute; ELLIS Unit Potsdam(苹果公司; 哈索·普拉特纳研究所; 波茨坦ELLIS单元)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对非英语及多语言场景开展GRPO大规模实证,发现母语训练推理与英语训练差距小、存在跨语言迁移但趋势依赖模型语言,指出非英语RLVR有跨语言增益但需全面评估退化问题。

AI 中文摘要

带可验证奖励的强化学习(RLVR)通常采用分组相对策略优化(GRPO)进行优化,已成为提升预训练语言模型推理能力的核心方法,但现有研究仍严重以英语为中心。我们针对多语言及非英语场景下的GRPO开展大规模实证研究,涉及多种基础模型、训练语言及不同推理语言奖励。研究发现,采用母语训练推理与采用英语训练推理仅存在微小差距;还观察到显著的跨语言迁移现象:在某一语言上训练常可提升其他多种语言的性能,但具体趋势高度依赖模型与语言,部分情况下在特定语言上训练会导致其他语言的域外能力严重退化。分析表明,非英语场景下的RLVR可带来广泛的跨语言增益,但也需要全面评估以检测特定语言的退化问题。

英文摘要

Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong crosslingual transfer: training in one language often improves performance in many others. However, specific trends are highly model- and language-dependent. In some cases, training in a particular language induces severe regressions on out-of-domain capabilities in other languages. Our analysis shows that RLVR beyond English can provide broad crosslingual gains, but also requires broad evaluation to detect language-specific regressions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑