非英语语言的推理成本:以日语为例
Cost of Reasoning in non-English Languages: A Case Study on Japanese
浏览论文内容
中文总结 AI 辅助
研究训练日语推理模型的可行性,开发Qwen - 3 - Swallow - 8B的日语推理变体并用GRPO训练,在多基准测试中发现推理语言控制可行,但性能与英语推理基线相当,在日本文化基准上表现不如基线模型。
中文摘要 AI 辅助
推理语言模型(RLMs)在使用英语推理时表现最强,因为英语有最丰富的面向推理的训练数据。然而,推理痕迹对模型的可解释性和安全性很重要,对用户和开发者都有用。因此,希望开发能在用户选择的语言中推理且保持强大推理性能的模型。为此,研究训练日语推理模型的可行性。开发了Qwen - 3 - Swallow - 8B的日语推理变体,用GRPO训练并在编码、数学和科学基准上评估。研究表明通过GRPO训练日语持续预训练模型实现推理语言控制是可行的,但在几个基准上其性能至多与强大的英语推理基线相当。在日本文化基准上评估时,该模型性能比基线模型差,表明日语推理不会自动提升与文化相关任务的性能。
英文摘要
Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training data is most abundant. However, reasoning trace is a clue for model interpretability and safety, and useful in practice for both the model users and for model developers. Thus, it is desirable to be able to develop a model that reasons in a language of the user's choice, while still maintaining strong reasoning performance. To this end, we study the feasibility of training a model that reasons in Japanese. We develop a Japanese-reasoning variant of Qwen-3-Swallow-8B, which is a Japanese LLM continually pretrained from Qwen-3-8B, with GRPO and evaluate it across coding, math, and science benchmarks. The study shows that reasoning-language control is feasible by training a Japanese continually pretrained model with GRPO. However, its performance is at best on par with strong English-reasoning baselines on several benchmarks. We also evaluate the trained model on Japanese cultural benchmarks and observe that the model's performance is worse than the baseline models, suggesting that the reasoning in Japanese does not immediately improve performance on culturally relevant tasks for free.