构建多语言桥梁:数据混合作为语言内推理泛化的支柱
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
- Cohere Labs(Cohere 实验室)
- Cohere
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出通过数据混合优化SFT,构建Tiny Aya L2-Thinker模型,在60种语言上实现超93%的L2推理率,证明推理可跨语言泛化。
AI中文摘要:
推理语言模型在各种复杂任务上取得了显著进展,但其能力仍以英语为中心:无论提示使用何种语言,模型主要用英语进行推理。这对非英语用户而言难以使用,有丢失原始问题意图的风险,并忽略了在目标语言中更易表达的知识。在本工作中,我们推进了L2推理,即模型在用户提示的语言中一致推理的能力,从而在提示与答案之间建立语言内桥梁。我们从数据为中心的角度处理这一问题,研究如何在SFT中优化数据组成与调度以实现推理泛化。我们构建了3.35B规模的Tiny Aya L2-Thinker,在涵盖数学、常识推理、指令遵循、开放式生成和文化推理的6个基准上,跨越60种语言实现了超过93%的L2推理率,同时保持强劲性能。我们展示了将L2推理泛化到未见语言需通过更广泛的语言覆盖、现成的多语言非推理数据以及足够的英语推理骨干来实现。这些发现表明,推理是一种与语言无关的行为,可通过仔细的数据混合在类型多样的语言间转移,而无需在每个目标语言中提供推理监督。我们发布模型权重和多语言推理数据,以支持关于可访问的语言内推理的进一步研究。
英文摘要:
Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: models primarily reason in English regardless of the language they are prompted in. This is inaccessible for non-English-speaking users, risks losing the intent of the original question, and forgoes knowledge more readily expressed in the target language. In this work, we advance L2 reasoning, the ability of a model to reason consistently in the language of the user's prompt, thus building an in-language bridge between the prompt and the answer. We approach this problem from a data-centric angle, investigating how to optimize data composition and scheduling in SFT for reasoning generalization. Building Tiny Aya L2-Thinker at 3.35B scale, we achieve an L2 reasoning rate above 93% across 60 languages on 6 benchmarks spanning math, commonsense reasoning, instruction following, open-ended generation, and cultural reasoning while keeping performance strong. We show the path to generalizing L2 reasoning to held-out languages goes through broader language coverage, readily available multilingual non-reasoning data, and a sufficient English reasoning backbone. These findings indicate that reasoning is a language-agnostic behavior that can be transferred across typologically diverse languages through careful data mixing and without requiring reasoning supervision in every target language. We release our model weights and multilingual reasoning data to support further research on accessible, in-language reasoning.