arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04980cs.CLcs.AIcs.LG

微型Transformer中的原型推理

Protoreasoning in Tiny Transformers

Eduardo Valle, Fergal Reid

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出原型推理,在约100万参数的微型Transformer上实现逐步推理,通过Dyck语言任务验证其可缩小分布外泛化差距,助力探究LLM的通用推理能力。

中文摘要 AI 辅助

我们证明,微型Transformer可以有效采用一种名为原型推理(protoreasoning)的简单思维链形式,这使我们能够在约100万参数的模型上研究逐步推理,并为比更大模型更详细的实验和分析创造机会。当前大型语言模型(LLM)展现出令人印象深刻的逐步推理能力,但我们尚未理解其通用性,即LLM何时以及如何学习真正通用的算法而非“启发式集合”。这类问题在基于不透明数据训练的计算密集型前沿模型上难以解决。为在远低于自然语言能力阈值的模型规模下开展工作,我们在Dyck语言(正确嵌套括号的句子)上定义了利于推理的任务。我们发现,原型推理大幅缩小了分布外泛化差距,且消融实验证实,该轨迹的内容而非仅额外的标记驱动了性能提升。

英文摘要

We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-step reasoning on ~1M-parameter models and opening up opportunities for much more detailed experimentation and analysis than is feasible for larger models. Current Large Language Models exhibit impressive step-by-step reasoning, but we have yet to understand its generality, i.e., when and how LLMs learn genuinely general algorithms rather than "bags of heuristics." Such questions are hard to settle on compute-intensive frontier models trained on opaque data. To work at model scales far below the threshold for natural-language competence, we define reasoning-friendly tasks on Dyck languages (sentences of correctly nested brackets). We find that protoreasoning traces substantially close the out-of-distribution generalization gap, and ablations confirm that the trace's content, not merely its extra tokens, drives the gain.

发表机构

  • Fin AI Research(Fin AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑