预雕刻的生态位:早期大语言模型训练中模块化任务分区的形成动力学
Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training
浏览论文内容
中文总结 AI 辅助
该研究追踪Pythia-410M模型训练中模块化任务分区的形成过程,发现其模块化图谱预雕刻、分区通过梯度相对剥夺锁定且偏差随学习领域出现,预注册了28亿参数实验的规模阈值假设。
中文摘要 AI 辅助
大型语言模型呈现出与已被充分研究的人类大脑功能网络相呼应的模块化内部组织,但这种组织如何在训练过程中形成尚不明确:现有研究仅分析了训练完成的模型,而非其形成过程。我们逐步追踪该形成过程:从零开始训练Pythia-410M模型(包含两条轨迹,分别为bf16和fp32精度),并在每一步执行归因修补,同时针对四个认知领域的14项任务,对梯度范数、有效更新量、权重范数以及一阶损失分解进行探测,得出三项发现。其一,模块化图谱是预雕刻的:在任何学习发生前,主导任务对已在归因基底(一项与任务无关的基线)上重叠约3.6倍,且其第0层集中度是该模型家族的架构级常数。其二,分区通过两次急剧跳跃锁定,其幅度不随学习率调度变化(第二次跳跃达到20.4倍安静窗口标准差/6.2倍全局标准差),同时伴随梯度层面的相对剥夺——获胜者获得的梯度供给为失败者的2.25至2.73倍,比随机对照低9.5至11.5个标准差,且该现象不会传播至更新量或权重。其三,与基底的偏差仅出现在正在学习的领域中,与模块化追踪学习的假设一致。最后,我们区分了可论证的特征层面解释与无法解答的机制性问题,并预先注册了支撑正在开展的28亿参数实验的规模阈值假设。
英文摘要
Large language models exhibit a modular internal organization that mirrors well-studied functional networks of the human brain, but how this organization forms during training is unknown: prior work has characterized finished models, not the formation process. We track formation step by step: we train a Pythia-410M model from scratch (two trajectories, bf16 and fp32) and run attribution patching at every step, alongside probes for gradient norms, effective updates, weight norms, and first-order loss decomposition across 14 tasks in four cognitive domains. Three findings. First, the modular map is pre-carved: before any learning, the dominant task pair already overlaps at ~3.6x the attribution substrate (a task-independent baseline), and its layer-0 concentration is an architecture-level constant on this model family. Second, the partition locks in through two sharp jumps whose amplitudes do not track the learning-rate schedule (the second reaching 20.4 sigma quiet-window / 6.2 sigma global), accompanied by gradient-level relative deprivation--winners receive 2.25->2.73x the loser's gradient supply, 9.5-11.5 standard deviations below a random control--that does not propagate to updates or weights. Third, deviation from the substrate appears only in the domain being learned, consistent with the hypothesis that modularity tracks learning. We close by separating the feature-level account we can defend from the mechanistic questions we cannot, and we pre-register the scale-threshold hypothesis behind our ongoing 2.8B experiments.
发表机构
- Zaozhuang University(枣庄学院)
机构由 AI 辅助整理,请以论文原文为准。