注意力作为路由图:单次前向传播的实时电路提取
Attention as a Routing Graph: Live Circuit Extraction from a Single Forward Pass
浏览论文内容
中文总结 AI 辅助
提出一种基于单次前向传播的注意力路由图方法,通过提取指向答案的少量边来近似语言模型中的电路,在已知任务上验证了其因果有效性,且成本远低于传统干预方法。
中文摘要 AI 辅助
在语言模型中寻找电路通常意味着进行许多细致的干预。我们尝试更简单的方法:将注意力视为单次前向传播中的路由图,保留指向答案的一小组路由,并询问这些路由是否真的重要。它们通常确实重要。在归纳和IOI任务(这些任务的“正确”电路已知)上,消融我们提取的边比消融相同大小的随机边对模型的损害要大得多。我们在GPT-2 Small、GPT-2 Medium和Pythia-410M上评估了每个单元n=100个提示,并进行了配对差距检验和bootstrap置信区间。提取步骤只需一次前向传播;逐头补丁扫描的成本大约高出两个数量级。我们并不声称提供完整的电路图谱。我们声称提供一种廉价的草图,它在已知任务上携带真实的因果信号,并在不适用时具有明确的失败模式。代码和评估工件位于此https URL。
英文摘要
Finding circuits in language models usually means running many careful interventions. We try something simpler: treat attention as a routing map from one forward pass, keep a small set of routes that point toward the answer, and ask whether those routes actually matter. They often do. On induction and IOI (tasks where the "right" circuit is already known), ablating our extracted edges hurts the model much more than ablating a random set of the same size. We evaluate n=100 prompts per cell on GPT-2 Small, GPT-2 Medium, and Pythia-410M, with paired gap tests and bootstrap confidence intervals. The extract step costs one forward; a head-by-head patch sweep costs about two orders of magnitude more. We are not claiming a complete circuit atlas. We are claiming a cheap sketch that carries real causal signal on known tasks, with clear failure modes when it does not. Code and evaluation artifacts are at https://github.com/Aquinf03/live-circuit-routing.
发表机构
- Aquin Labs(Aquin 实验室)
机构由 AI 辅助整理,请以论文原文为准。