不太难也不太易:从中间状态学习以进行LLM结构化推理
Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning
中文总结 AI 辅助
本研究提出FOCUS方法,通过从模型自生成轨迹中筛选处于能力前沿的中间状态进行训练,使语言模型在数独和迷宫等结构化推理任务上显著提升求解准确率,并展现零样本迁移能力。
中文摘要 AI 辅助
有效学习的一个常见原则是练习那些既未完全掌握又不至于难到无法取得进展的材料。我们探讨如何将此原则应用于诸如数独和迷宫求解等结构化推理任务。在这些任务中,模型可以反复修改不完整或错误的候选解,直到其满足问题的约束条件。沿此轨迹产生的中间候选解提供了自然的训练样本:有些已解决,有些模型尚无法修复,而另一些则处于模型当前可取得进展的前沿。因此,我们研究预训练语言模型能否学会修改此类状态,以及在此前沿状态上训练是否能在更广泛范围内提升推理能力。为实现此目标,我们将预训练语言模型主干与一个循环更新器相结合,该更新器使用相同的参数在每一步更新中反复修改显式解状态。我们进一步引入了基于自轨迹的前沿导向筛选(FOCUS),该方法从当前模型生成的轨迹中选择训练状态。FOCUS衡量模型在固定次数的循环更新内对每个状态的改进程度,并优先选择那些模型能取得实质性进展的状态。使用Qwen3-1.7B,FOCUS在Sudoku-Extreme上达到64.4%的精确求解准确率,在Maze-Hard上达到91.1%,在跨越1.7B至8B参数的五个Qwen和Llama主干上观察到类似增益。我们还观察到,即使禁用循环更新器且不进行下游微调,适配后的LLM也能在数学推理和代码执行方面实现零样本迁移。
英文摘要
A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze solving. In these tasks, a model can repeatedly revise an incomplete or incorrect candidate solution until it satisfies the problem's constraints. The intermediate candidate solutions along this trajectory provide natural training examples: some are already solved, some cannot yet be repaired by the model, and others lie at its current frontier of achievable progress. We therefore investigate whether pretrained language models can learn to revise such states and whether training on states at this frontier improves reasoning more broadly. To achieve this, we couple a pretrained language-model backbone with a recurrent updater that repeatedly revises an explicit solution state, using the same parameters at every update step. We further introduce Frontier-Oriented Curation Using Self-trajectories (FOCUS), which selects training states from trajectories generated by the current model. FOCUS measures how much the model improves each state within a fixed number of recurrent updates and prioritizes states from which it can make substantial progress. With Qwen3-1.7B, FOCUS achieves 64.4% exact solve accuracy on Sudoku-Extreme and 91.1% on Maze-Hard, with similar gains observed across five Qwen and Llama backbones spanning 1.7B to 8B parameters. We further observe zero-shot transfer in the adapted LLM to mathematical reasoning and code execution, even when the recurrent updater is disabled and no downstream fine-tuning is performed.