AI 中文总结
研究针对前馈网络占Transformer多数非嵌入参数却缺相关必要性测试的情况,预训练仅注意力机制的解码器Transformer并与标准Transformer对照,发现删除前馈层成本高,重分配预算可缩小差距,还定位了剩余差距原因,预注册测试证实相关解释。
AI 中文摘要
前馈网络占据了Transformer三分之二的非嵌入参数,但该架构尚未接受同时控制参数、计算量和深度的必要性测试。我们针对参数数量、训练FLOP和深度(2至48层)分别匹配的标准Transformer,预训练仅注意力机制的解码器Transformer(简单注意力网络,SANs),参数规模从6M到87M,训练token数可达105B。直接删除前馈层成本高昂:在匹配深度时,标准Transformer领先0.47奈特,在匹配计算量时领先0.26奈特。将节省的预算重新分配到注意力深度可缩小差距:在匹配参数时,差距为0.006奈特(损失的0.27%),在不同种子对中可重复到万分之一,在5B、30B和105B预算下差距缩小,在29倍规模范围内接近0.02奈特。三项测量将剩余差距定位到参数召回:仅注意力机制模型在基于上下文的答案上表现更好,在知识必须来自权重的地方表现更差。权重谱显示了原因:路由矩阵(Q/K)早期结晶,内容矩阵秩积累缓慢,删除前馈层将这种积累转移到注意力输出投影。QK归一化而非前馈层或残差门控使48层仅注意力机制堆栈可训练。差距集中在低上下文查询预测上,并在最大预算下完全定位在此。预注册测试证实了这一解释:它预测在知识密集的网络文本上差距为0.02至0.05奈特;在fineweb-edu上训练的匹配对测量值为0.040。在测试范围内,其余由注意力完成。
英文摘要
Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test that controls parameters, compute, and depth at once. We pretrain attention-only decoder transformers (Simple Attention Networks, SANs) against standard transformers matched separately for parameter count, training FLOPs, and depth (2 to 48 layers), for up to 105B tokens at 6M to 87M parameters. Deleting feed-forward layers in place is costly: the standard transformer leads by 0.47 nats at matched depth and 0.26 nats at matched FLOPs. Reallocating the freed budget into attention depth closes the gap: at matched parameters the difference is 0.006 nats (0.27 percent of loss), reproducible to one part in ten thousand across seed pairs, shrinking across 5B, 30B, and 105B budgets, and holding near 0.02 nats across a 29x size range. Three measurements localize the remaining gap to parametric recall: attention-only models are better on context-grounded answers and worse where knowledge must come from weights. Weight spectra show why: routing matrices (Q/K) crystallize early, content matrices accumulate rank slowly, and removing feed-forward layers relocates this accumulation to the attention output projection. QK-normalization, not feed-forward layers or residual gating, keeps 48-layer attention-only stacks trainable. The deficit concentrates on low-context query prediction and localizes there entirely by the largest budget. A pre-registered test confirms the account: it predicts a 0.02 to 0.05 nat gap on knowledge-dense web text; a matched pair trained on fineweb-edu measures 0.040. Within the tested regime, attention does the rest.
Comments10 pages, 8 figures