发表机构
School of Engineering, Institute of Science Tokyo; College of Control Science and Engineering, Zhejiang University; Department of Electrical and Computer Engineering, National University of Singapore(科学东京大学工程学院; 浙江大学控制科学与工程学院; 新加坡国立大学电气与计算机工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现多语言RLVR中存在验证器偏差,引入审计协议定位到最终答案接口的机制,提出跨语言选择瓶颈,rule-GRPO可提升可信准确率。
AI 中文摘要
带可验证奖励的强化学习(RLVR)是在数学推理任务上训练大语言模型的标准方法,其中答案验证器作为与语言无关的奖励函数。本文表明该假设在多语言场景下不成立:精确匹配验证器会将格式和脚本差异转化为依赖语言的假负奖励噪声。我们引入了可复用的多语言RLVR奖励审计协议,包括验证器鲁棒性套件、推理诊断流程,以及针对日语、英语和中文答案的语言条件奖励误差指标。在MGSM的k=8推理中,Qwen3-4B、Qwen3-8B和Llama-3.1-8B-Instruct模型的精确匹配代理以显著不同的速率拒绝可信正确答案;对于Qwen3-8B,日语上的假负率达0.642,而英语为0.122,中文为0.073。纯数值探针将该机制定位到最终答案接口:接口模型使奖励误差VLB归零,同时剩余准确率差距保持不变。我们进一步揭示了跨语言选择瓶颈:在MGSM250推理中,不使用可信标签的目标局部聚合规则可缩小55%-78%的平均选择差距,且超过95%的修复需要真正的跨语言支持;该瓶颈在含483个问题的MATH-500数据集上可复现。受控训练审计显示,rule-GRPO可提升可信准确率,同时奖励误差VLB保持高位。核心结论为:多语言RLVR奖励在优化前应按语言和答案接口进行审计。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) is a standard recipe for training large language models on mathematical reasoning, where an answer verifier serves as a language-neutral reward function. We show that this assumption fails in multilingual settings: an exact-match verifier turns format and script variation into language-dependent false-negative reward noise. We introduce a reusable protocol for auditing multilingual RLVR rewards: a verifier-robustness suite, a rollout-diagnosis procedure, and language-conditioned reward-error metrics for Japanese, English, and Chinese answers. On MGSM rollouts with k=8, the exact-match proxy rejects trusted-correct answers at sharply different rates by language across Qwen3-4B, Qwen3-8B, and Llama-3.1-8B-Instruct; for Qwen3-8B, the false-negative rate reaches 0.642 on JP against 0.122 on EN and 0.073 on CN. A plain-numeric probe localizes the mechanism to the final-answer interface: an interface model drives reward-error VLB to zero while the residual accuracy gap is unchanged. We then expose a cross-lingual selection bottleneck: on MGSM250 rollouts, a target-local aggregation rule using no trusted labels closes 55-78% of the average selection gap, and over 95% of repairs require genuine cross-lingual support. The bottleneck replicates on a 483-problem MATH-500 set. A controlled training audit shows that rule-GRPO raises trusted accuracy while the reward-error VLB stays high. The unifying message is operational: multilingual RLVR rewards should be audited by language and by answer interface before they are optimized.
Comments16 pages, 2 figures, 5 tables