arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

每一次消融都是一次剂量:配重与自我修复的表象

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu

arXiv 2610.02173首次发表:更新:

发表机构

Lexsi Labs(Lexsi 实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出消融可视为剂量,揭示语言模型自我修复实为配重效应,通过仿射定律预测修复响应,并在多个模型上验证。

AI 中文摘要

消融语言模型的一个组件,其他组件往往看起来会进行调整和补偿。这种现象被称为自我修复,已被反复观察到,但其机制仍不清楚。迄今为止最系统的研究得出结论,自我修复是嘈杂的,不太可能有一个单一的解释。我们认为它有一个:在消融之前就已经存在的一个增益。对因果重要组件的任何干预都可以被视为坐标轴 λ 上的一个点,即反事实对比的带符号强度。因此,传统的消融方法是该轴上未校准的点。我们表明,对于细粒度单元 r,因果修复响应由仿射定律 E_r(λ)=own_r+γ_rλ 控制。斜率 γ_r 是一个固定系数,无论是否消融,它都持续影响模型,其符号决定该单元是抵消还是增强被移除的信号。在一个事实判定任务上,跨越四个不同家族的模型(Gemma、Qwen、LLaMA 和 Mistral),我们识别出包括 MLP 神经元、OV 神经元和奇异方向在内的组件遵循该仿射定律,总共 81 个下游方向中有 68 个。此外,我们可以从固定权重中预期 γ_r 的大小。在 GPT-2 Small 的 IOI 电路中,干预可及的十个头中有七个遵循该定律,且全部七个都是配重。从这个角度来看,当对比信号在核心出现时,看似自我修复的现象实际上是配重执行其常规操作。

英文摘要

Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $γ_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑