arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17744cs.CLcs.LGcs.ROstat.ML

在低资源语言中思考:SFT构建了什么,RL修正了什么,以及准确率无法察觉的问题

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

Ayoub Kirouane, Christos Petrocheilos

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对三个混合专家模型,发现低资源语言推理中SFT可提升推理语言一致性与流畅性,RL可修正格式遗漏与答案泄露问题,而仅靠准确率无法察觉关键变化,相关工具可推广至其他低资源语言。

中文摘要 AI 辅助

选取三个前沿的混合专家模型(Alibaba、OpenAI、NVIDIA,每个模型的激活参数规模为36亿至40亿),对它们进行微调以实现低资源语言推理。在准确率基准测试中几乎没有任何变化,而且在该规模下,基准测试本身存在噪声:仅改变随机种子就会使分数波动7.7个点,超过了我们测量到的所有数据和训练方案的影响,这一无效结果是我们的首个发现。真正的变化发生在准确率无法察觉的领域:基础模型永远不会用希腊语思考——在1000条推理轨迹中,即使问题是希腊语,也没有一条是用希腊语推理的,因此模型在正确回答问题的同时,会以用户无法阅读、审计或修正的形式进行推理。经过监督微调(SFT)后,所有发布的检查点在约98%的项目中都能使用问题的语言进行推理,其中一个系列的推理使用的token数减少了3倍,经评估的语法正确性在所有四个模型上均有所提升,且通用能力与基础模型仅相差几个点:没有遗忘任何能力,同时获得了流畅性。我们提出了六个行为维度来使这类变化可测量,每个维度都设置了门限以拒绝任何与输出长度相关的指标,并且我们报告了自己的工具存在的六个故障,每个故障都通过对照实验被发现。SFT无法修正自身的缺陷:四分之一的答案会跳过要求的格式,答案会泄露到推理通道,且明确要求“用英语思考”的指令仅有不到一半的时间被遵守。使用可验证奖励的强化学习(RL)在训练前已预先注册,它彻底修正了前两个问题(格式遗漏从24%降至2.5%,泄露从3.5%降至0.0%,均与随机奖励对照实验相比),并使第三个问题的表现提升了9.1个百分点,而仅基于准确率的梯度训练无法改变希腊语推理的习惯。我们发布了五个检查点。这些工具、对照实验和预先注册的方法可应用于任何低资源语言;希腊语是让我们能够测量这些变化的案例。

英文摘要

Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.

发表机构

  • Sophea AI(索非亚人工智能公司)
  • KIEFER SA(基弗股份有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑