arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03620cs.LGcs.AI

低秩权重空间消融下的条件坍缩理论:I. 单块理论与合成验证

A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation

Abdallah Khemais

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出低秩权重空间消融下的条件坍缩单块理论,通过合成任务实验验证其预测,揭示激活修补与权重空间消融的差异及交互规律。

中文摘要 AI 辅助

激活修补(activation patching)和权重空间消融(weight-space ablation)均声称某一组件对特定行为具有因果责任,但二者作用对象不同:前者作用于单次前向传播,后者作用于每次前向传播背后的参数。本文探究二者何时会达成一致。我们研究一种理想化模型,其中条件计算通过残差流加性执行,即$F(x)=F_0(x)+\sum_i\alpha_i(x)v_i$,并通过线性泛函读出,据此证明三个精确结论。第一,删除一部分载体将匹配输入对坍缩至相同无条件输出,当且仅当该删除对输入对对称且不留下外部对比;误差是确定性的,且即使两个条件仅近似成立,我们也给出其精确形式。第二,修补载体将读出量移动其供体-受体对比,而消融载体将读出量移动其绝对水平;二者互不约束,我们构造出这样的输入对:每个单载体修补都会翻转决策,而无任何单载体消融会如此。第三,对于由自身层的归一化和多层感知机(MLP)组成的注意力头,我们推导了一个精确的一阶交互公式,其可证二阶余项在仅消融MLP时恒为零,但在消融注意力头时一般不为零。在合成条件任务上训练的小型Transformer验证了所有三个预测:在39种消融配置中,测得的交互作用与理想化模型的预测准确率具有强秩相关性(Spearman秩相关系数为$-0.83$),且另一任务和架构也复现了相同模式,包括进一步的极性反转。该单块交互结果可扩展至一个以上残差块,合成验证还在一篇姊妹论文中针对真实预训练模型进行了测试,该姊妹论文沿两个方向进一步发展了此理论。

英文摘要

Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass. We ask when they agree. We study an idealized model where a conditional computation is carried additively through a residual stream, $F(x)=F_0(x)+\sum_iα_i(x)v_i$, read out by a linear functional, and prove three exact results. First, deleting a subset of carriers collapses a matched input pair onto the same unconditional output \emph{if and only if} the removal is symmetric on the pair and leaves no outside contrast; the error is deterministic, and we give its exact form even when the two conditions hold only approximately. Second, patching a carrier moves the readout by its donor-receiver \emph{contrast}, while ablating it moves the readout by its \emph{absolute level}; neither bounds the other, and we construct pairs where every single-carrier patch flips the decision while no single-carrier ablation does. Third, for an attention head composed with its own layer's normalization and MLP, we derive an exact first-order interaction formula with a provably second-order remainder, vanishing identically when only the MLP is ablated but not, in general, when a head is. Small transformers trained on a synthetic conditional task illustrate all three predictions: across thirty-nine ablation configurations the measured interaction is strongly rank-correlated with the idealized model's predictive accuracy (Spearman $-0.83$), and a second task and architecture reproduces the same pattern, including a further polarity reversal. The single-block interaction result extends past one residual block, and the synthetic validation is tested against a real pretrained model, in a companion paper that takes this theory further along both axes.

补充信息

↑