发表机构
University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出模式识别与逐步推理是谱系两端,通过De Bruijn图建模推理轨迹,证明边数远少于轨迹数,并实证表明适度状态密度可平衡准确性与鲁棒性,且真实数据具有De Bruijn结构。
AI 中文摘要
我们认为模式识别与逐步推理是谱系的两端。当数据结构使得下一个词元仅依赖于少量先前上下文时,大型语言模型(LLM)学会逐步推理。当下一个词元依赖于大量先前上下文时,LLM中的推理类似于模式识别。如果下一个词元仅依赖于最近的$c$个词元,则推理轨迹是De Bruijn图上的路径,其节点为$c$长度的上下文,边为上下文之间的下一个词元转移。任务的一组推理轨迹构成De Bruijn图的一个有向无环子图。已学习该子图所有边的LLM能够组合这些边以解决更长、未见过的任务,即逐步推理。我们证明边的数量与推理轨迹数量相比是极小的。实证上,transformer所需的训练样本数量是边数的幂律,因此逐步推理学习是样本高效的。我们可通过维护一个“状态”来在任何任务中诱导De Bruijn结构,该状态使未来推理独立于过去。推理轨迹中状态的频率决定$c$。通过微调Qwen2.5-1.5B-Instruct以解方程和回答故事相关问题,我们证明频繁状态(小$c$)导致更高准确率但在测试时对扰动更脆弱。以较大$c$训练的LLM仅与执行无推理的模式识别的模型相当。适度的状态密度平衡了准确性与鲁棒性。我们表明真实世界数据具有De Bruijn结构:当注意力限制在比完整推理轨迹短不到15%的滑动窗口时,Qwen3-14B和Qwen3-32B在GSM8K、MATH-500和GPQA-Diamond上保持超过75%的准确率。
英文摘要
We argue that pattern recognition and step-by-step reasoning are two ends of a spectrum. A large language model (LLM) learns to reason step-by-step when data is structured such that the next token depends on a small amount of preceding context. Inference in LLMs resembles pattern recognition when the next token depends on a large amount of preceding context. If the next token depends on only the $c$ most recent tokens, reasoning traces are paths on a De Bruijn graph whose nodes are $c$-length contexts and edges are next-token transitions between contexts. The set of reasoning traces of a task forms a directed acyclic subgraph of the De Bruijn graph. An LLM that has learned all edges of this subgraph can compose them to solve longer, unseen tasks, i.e., it reasons step-by-step. We prove that the number of edges is vanishingly small compared to the number of reasoning traces. Empirically, the number of training samples a transformer needs is a power law in the number of edges, so learning to reason step-by-step is sample efficient. We can induce De Bruijn structure in any task by maintaining a ``state'' that makes future reasoning independent of the past. The frequency of states in the reasoning trace determines $c$. We show, by fine-tuning Qwen2.5-1.5B-Instruct to solve equations and answer questions about stories, that frequent states (small $c$) result in higher accuracy but greater fragility to perturbations at test time. LLMs trained with a large $c$ are only as good as models that perform pattern recognition without reasoning. A moderate density of states balances accuracy and robustness. We show that real-world data has De Bruijn structure: Qwen3-14B and Qwen3-32B retain over 75% of their accuracy on GSM8K, MATH-500 and GPQA-Diamond when attention is restricted to a sliding window less than 15% as long as the full reasoning trace.