预LN Transformer中部分残差消融的可复现性研究
A Reproducibility Study of Partial Residual Ablations in Pre-LN Transformers
浏览论文内容
中文总结 AI 辅助
本文通过对10M和124M参数的预LN GPT风格Transformer开展部分残差消融可复现性研究,揭示注意力与前馈残差路径的不同作用,修正实验混淆并发布完整实验材料。
中文摘要 AI 辅助
残差连接是Transformer架构的基础组件,但注意力残差路径与前馈残差路径的独立作用仍未被充分理解。本文对两种规模(10M和124M参数)的预LN GPT风格Transformer开展部分残差消融的可复现性研究,通过选择性移除注意力残差连接、前馈残差连接或两者,对比四种架构配置。所有实验中,移除注意力残差(FFNOnly)会一致导致确定性崩溃至无残差性能下限;而移除前馈残差(AttnOnly)在10M规模下,经8个种子的确定性受控研究呈现可复现的恢复效应,但其在124M规模下的行为因种子方差较大仍未明确。研究期间,作者识别并修正了运行时增益缩放中的实验测量混淆,记录了失败的中间复现及后续受控复现过程;基于实证结果,提出跨位置路由假设以解释观察到的不对称性,并明确区分已确认发现与未解决问题。为支持可复现性,作者发布了完整源代码、实验配置、检查点、训练日志及所有实验结果,包括中间未复现的运行。
英文摘要
Residual connections are a fundamental component of transformer architectures, yet the roles of the attention and feed-forward residual pathways remain poorly understood when considered independently. This paper presents a reproducibility study of partial residual ablations in Pre-LN GPT-style transformers trained at two scales (10M and 124M parameters). I compare four architectural configurations by selectively removing the attention residual connection, the feed-forward residual connection, or both. At 10M scale, a controlled 8-seed deterministic sweep shows a clear asymmetry: removing the attention residual (FFNOnly) reaches the No-Residual collapse regime (3.350 +/- 0.002), whereas removing the feed-forward residual (AttnOnly) remains well separated from it (1.580 +/- 0.003). A subsequent controlled 124M sweep establishes that the direction of this asymmetry persists under matched seeds: across five seeds each, AttnOnly reaches 6.037 +/- 0.355 versus 7.569 +/- 0.001 for FFNOnly, with no overlap between the observed ranges. AttnOnly nevertheless remains substantially degraded relative to FullResidual (4.577 +/- 0.012) and exhibits markedly greater seed sensitivity at 124M. During the investigation, I identified and corrected an experimental measurement confound in runtime gain scaling and retained an intermediate reproduction failure rather than excluding it. I propose cross-position routing through self-attention as a falsifiable hypothesis for the observed asymmetry; no clean causal test of this mechanism has yet been completed. To support reproducibility, I release the source code, experiment configurations, training logs, and experimental results, including intermediate non-reproducing runs.