arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33559cs.LG

直到微调前都很好:重复解决方案使推理变得脆弱

Fine Until Fine-Tuned: Repeated Solutions Make Reasoning Fragile

  • Aix-Marseille University(艾克斯-马赛大学)

机构由 AI 辅助整理,请以论文原文为准。

Ely Sheikh

AI总结:

重复展示相同解决方案的推理训练虽在当下无害,却使模型在后续微调中推理能力显著下降,新解决方案或重放部分数据可避免此脆弱性。

AI中文摘要:

诸如s1和LIMO之类的配方通过多次向模型展示相同的一千个或更少的已解答解决方案,用很少的数据教会模型进行推理。在训练结束时进行评判,这种重复看起来无害。但推理模型通常会被再次训练,我们发现重复会使它们的推理对下一阶段变得脆弱,即使该阶段与推理无关。我们在Qwen3.5-9B-Base上使用其自身对竞赛数学问题的正确解决方案进行微调,要么对几百个解决方案各钻约八次,要么展示更多不同的解决方案各一次;在相同训练量下,两者都能解决约95%的留出问题。一次普通的指令微调使仅训练一次的模型保持原状,而经过钻探的模型下降到86.0%,更严苛的后续阶段将其降至59.3%或更低。第三个模型以相同的频率访问钻探过的问题,但每次访问都使用新的解决方案,结果未受损害,因此损害来自重复看到相同的文本,而非问题数量少。这种破坏在更强的模型轨迹、进一步的训练运行以及其他模型和任务中反复出现。它也很容易消除:推理被抑制而非抹除,五次推理训练更新几乎能完全恢复,对推理格式进行短暂训练且几乎不含数学也能恢复。新解决方案防止了损害,在较温和的后续阶段重放6.25%的原始解决方案也能防止,因此我们的主张涉及没有此类重放的后续训练。仅锐化并不能解释这种破坏,因为一个未经重复而锐化四分之三的模型未受损害。在基础模型无法在令牌预算内完成的技能上,重复主要代价是学习。

英文摘要:

Recipes such as s1 and LIMO teach a model to reason with little data by showing it the same thousand or fewer worked solutions many times over. Judged when that training ends, the repetition looks harmless. But reasoning models are often trained again, and we find that repetition leaves their reasoning fragile to that next stage, even when the stage has nothing to do with reasoning. We fine-tuned Qwen3.5-9B-Base on its own correct solutions to competition math problems, either drilling a few hundred of them about eight times each or showing many more once; with the same amount of training, both solve about 95% of held-out problems. A single pass of ordinary instruction tuning leaves the once-trained model where it was, while the drilled one falls to 86.0%, and harsher later stages take it to 59.3% or below. A third model that visited the drilled problems just as often, with a new solution at every visit, was unharmed, so the damage comes from seeing the same texts again rather than from having few problems. The break recurs with a stronger model's traces, in further training runs and on other models and tasks. It is also cheap to undo: the reasoning is suppressed rather than erased, and five updates of reasoning training bring almost all of it back, as does brief training on the reasoning format with almost no mathematics. Fresh solutions prevented the damage, and so did replaying 6.25% of the original solutions in a gentler later stage, so our claim concerns later training without such replay. Sharpening alone does not explain the break, since a model sharpened three-quarters as much without repetition was unharmed. On a skill the base model could not perform within a token budget, repetition mainly cost learning.

补充信息

↑