arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从神经网络中定向恢复权重空间机制

Targeted Recovery of Weight-Space Mechanisms From Neural Networks

Antoine Vigouroux, Lee Sharkey

arXiv 2607.13047首次发表:更新:

AI 中文总结

研究提出定向参数分解(tPD)方法,通过引入高秩通用组件,从神经网络中仅识别处理特定感兴趣输入的组件,在玩具模型和Transformer语言模型上验证有效,能以低计算量提取子模型并对记忆序列进行手术式操作。

AI 中文摘要

参数分解(PD)将神经网络分解为可解释的计算组件,能如实反映原始网络的操作。然而,将PD扩展到大型模型需要大量计算,成本高且风险大。本文提出定向PD(tPD),通过引入一个处理所有非目标数据的高秩通用组件,仅识别处理特定感兴趣输入(从孤立提示到大型子任务)的组件。我们在玩具模型和基于The Pile训练的Transformer语言模型上验证了tPD,它能恢复可重现、机制上忠实的电路。我们用其已发表分解7%的FLOP提取了一个4层Transformer的仅CSS子模型,在一个12层Transformer中,我们对记忆序列进行了手术式消融和重新布线,对其他输入的副作用可忽略不计。

英文摘要

Parameter decomposition (PD) decomposes neural networks into interpretable computational components that faithfully reflect the original network's operations. However, scaling PD to large models requires vast compute, making it a costly and risky endeavor. Here we propose targeted PD (tPD), which identifies only the components that process specific inputs of interest -- from isolated prompts to large subtasks -- by introducing a high-rank catch-all component that handles all non-target data. We validate tPD on toy models and on transformer language models trained on The Pile, where it recovers reproducible, mechanistically faithful circuits. We extract a CSS-only submodel of a 4-block transformer using 7% of the FLOPs of its published decomposition, and in a 12-block transformer we surgically ablate and rewire memorized sequences, with negligible side effects on other inputs.

CommentsAccepted at the Mechanistic Interpretability Workshop, ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑