发表机构
Beijing University of Civil Engineering and Architecture; Beijing Institute of Mathematical Sciences and Applications (BIMSA); Institute of Statistics and Big Data, Renmin University of China(北京建筑大学; 北京雁栖湖应用数学研究院; 中国人民大学统计与大数据研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过追踪预训练过程中不同检查点的块旁路响应,发现深度排序保持稳定但幅度变化,且变化集中在重复位置,表明层敏感性有结构但非静态,单检查点干预需结合训练演化背景。
AI 中文摘要
层干预被广泛用于探究语言模型的内部组织,然而大多数分析仅检查单个训练检查点,尽管模型表示和计算在整个预训练过程中不断演化。这留下了哪些依赖于深度的干预响应反映了持久组织,哪些是训练的瞬时后果的问题。我们使用固定教师强制上下文中的单块身份旁路,在五个发布的轨迹和11个模型-领域组合上研究这一问题。我们发现块旁路响应保持可识别的深度排序,而其幅度重新分布:相邻检查点比远处检查点保留更强的秩对应关系,且大的变化集中在跨文本样本重复出现并跨评估领域转移的位置。受控实验进一步表明,自然旁路效应的变化不能简化为单一的下游敏感性:在复制的Pythia运行中,局部缺失更新幅度增长,而池化匹配的下游响应下降,而OLMo-2 7B表现出不同的平衡。这些匹配响应也依赖于扰动强度和方向,而未识别出针对性的补偿。综合来看,我们的结果表明纵向层敏感性是有结构的但并非静态的,并且单检查点干预响应应在底层扰动路径在训练过程中如何演化的背景下进行解释。
英文摘要
Layer interventions are widely used to probe the internal organization of language models, yet most analyses examine a single training checkpoint even though model representations and computations evolve throughout pretraining. This leaves open which depth-dependent intervention responses reflect persistent organization and which are transient consequences of training. We study this question using single-block identity bypass on fixed teacher-forced contexts across five released trajectories and 11 model-domain combinations. We find that block-bypass responses retain recognizable depth ordering while their magnitudes redistribute: nearby checkpoints preserve stronger rank correspondence than distant ones, and large changes concentrate at positions that recur across text samples and transfer across evaluation domains. Controlled experiments further show that changes in the natural bypass effect cannot be reduced to a single downstream sensitivity: in replicated Pythia runs, local missing-update magnitude grows while the pooled matched downstream response decreases, whereas OLMo-2 7B exhibits a different balance. These matched responses also depend on perturbation strength and direction, without identifying targeted compensation. Together, our results show that longitudinal layer sensitivity is structured but not static, and that single-checkpoint intervention responses should be interpreted in the context of how the underlying perturbation pathway evolves during training.
Comments24 pages, 12 figures