算法草稿纸与课程分阶段在微型Transformer中的算术推理
Algorithmic Scratchpads and Curriculum Staging for Arithmetic Reasoning in Tiny Transformers
浏览论文内容
中文总结 AI 辅助
本研究在微型Transformer上探索多步算术推理,发现序列填充、预训练和架构组件的影响,提出逐位长除法草稿纸与课程分阶段使除法准确率大幅提升,并揭示乘法草稿纸设计缺陷及泛化与遗忘边界。
中文摘要 AI 辅助
自回归大型语言模型(LLMs)在多步确定性算法任务(如多位数乘法和长除法)上经常表现不佳。本文研究了紧凑型“微型”Transformer(约1060万非嵌入参数,总计4930万参数)在多步算术中的机制,这些模型在合成数据上训练,涵盖四种基本运算(+、-、*、/),并以逐步草稿纸的形式展开。首先,我们确立了必要的训练基础:(1)数据加载器序列填充造成83%的梯度饥饿伪影,使准确率从40%骤降至1%,通过连续序列打包得以修复;(2)语言预训练是必要前提(没有它准确率≤2.0%);(3)现代架构原语(RoPE、RMSNorm、SwiGLU)和稀疏专家混合(MoE)在加法推理上显著优于基线GPT-2。其次,我们证明算法草稿纸的表述直接决定成败。引入确定性的逐位长除法草稿纸,并在4阶段层次化发展课程中,将单位数除法在4000个问题的保留基准上的准确率从4.0%提升至86.7%。相比之下,多位数乘法仍然具有挑战性:详细的错误分析显示,虽然模型正确计算了单位数子乘积和位值零,但我们的FOIL草稿纸失败,因为它强制在单步中同时求和多达九个多位数项,而没有成对的中间累积。最后,我们识别出两个关键边界:在未见过的4位数操作数上性能骤降至0.00%,而无缓冲训练引发灾难性遗忘,使除法准确率从86.7%降至0.00%。
英文摘要
Autoregressive Large Language Models (LLMs) frequently struggle with deterministic multi-step algorithmic tasks such as multi-digit multiplication and long division. In this paper, we investigate the mechanics of multi-step arithmetic in compact "Tiny" Transformers (~10.6M non-embedding parameters, 49.3M total) trained on synthetic data across four basic operations (+, -, *, /) unrolled as step-by-step scratchpads. First, we establish the necessary training foundations: (1) dataloader sequence padding creates an 83% gradient starvation artifact that collapses accuracy from 40% to 1%, remediated via continuous sequence packing; (2) linguistic pretraining is an essential prerequisite (<= 2.0% without it); and (3) modern architectural primitives (RoPE, RMSNorm, SwiGLU) and Sparse Mixture of Experts (MoE) substantially improve additive reasoning over baseline GPT-2. Second, we demonstrate that algorithmic scratchpad formulation directly dictates success. Introducing a deterministic Digit-by-Digit Long Division scratchpad within a 4-stage Hierarchical Developmental Curriculum dramatically elevates single-digit division from 4.0% to 86.7% accuracy on a 4,000-problem held-out benchmark. In contrast, multi-digit multiplication remained challenging: detailed error analysis revealed that while the model correctly computed single-digit sub-products and place-value zeros, our FOIL scratchpad failed because it forced a simultaneous summation of up to nine multi-digit terms in a single step without pairwise intermediate accumulation. Finally, we identify two key boundaries: performance collapses to 0.00% on unseen 4-digit operands, and unbuffered training induces catastrophic forgetting, collapsing division accuracy from 86.7% down to 0.00%.