arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

电路声明取决于提取的内容以及比较的方式

Circuit Claims Depend on What Is Extracted and How It Is Compared

Yang Sheng, Jie Fu

arXiv 2607.18921首次发表:更新:

发表机构

Fudan University; Shanghai Innovation Institute; IQuest Research(复旦大学; 上海创新研究院; IQuest研究公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究指出电路提取中保留行为不能唯一确定电路,其声明依赖提取及比较方式。通过合成基准测试,在不同检查点和证明类型下改变提取对象等设置,发现边重叠敏感,总结有稳定指标,强调明确提取及比较要素对电路声明的重要性并提炼报告实践。

AI 中文摘要

电路提取识别出一小部分模型组件,其存在会在消融时保留目标行为,所得电路常被视为该行为背后的机制。我们认为这种解读是不充分的:保留行为并不能唯一确定一个电路,因为所支持的声明取决于报告的是哪个电路以及如何比较两个电路。我们在一个合成的精益策略预测基准测试中对此进行了具体说明——预测证明的下一步——其中具有随机表面形式的固定证明规则使提取电路之间的差异可归因于这些选择而非任务。在同一变压器的密集和权重稀疏检查点(大多数权重被约束为零)上,针对原子(单规则)和组合(多规则)证明进行评估,我们改变报告的提取对象(一个紧凑的保留预测的电路、一个更宽泛的图,它还保留周围的读、写和路由结构,或者满足消融后损失阈值的最小子图),以及每个注意力头的查询和键是联合表示还是单独表示。精确的组件到组件的边重叠率很低且对这些选择敏感,有时会降至随机基线,而两个更粗略的总结保持稳定:所选注意力头的集合,以及在哪个监督检查点初始化强化学习(RL)不同的条件下的电路大小排名。RL在组合证明上的最大准确率提升来自原子电路之外的最多结构。因此,只有在说明报告的是哪个电路、用于提取它的修剪阈值以及比较电路的级别时,电路级声明才是明确的。我们将这些要求提炼成电路提取研究的报告实践。

英文摘要

Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting circuit is often read as the mechanism behind that behavior. We argue that this reading is under-determined: preserving behavior does not single out one circuit, because the claim it supports depends on which circuit is reported and how two circuits are compared. We make this concrete in a synthetic Lean tactic-prediction benchmark -- predicting the next step of a proof -- where fixed proof rules with randomized surface form let differences between extracted circuits be attributed to these choices rather than to the task. Across dense and weight-sparse checkpoints (most weights constrained to zero) of the same transformer, evaluated on atomic (single-rule) and compositional (multi-rule) proofs, we vary which extracted object is reported (a compact prediction-preserving circuit, a broader graph that also keeps surrounding read, write, and routing structure, or the smallest subgraph meeting a post-ablation loss threshold), and whether each attention head's query and key are represented jointly or separately. Exact component-to-component edge overlap is low and sensitive to these choices, at times dropping to a random baseline, while two coarser summaries stay stable: the set of selected attention heads, and the circuit-size ranking of conditions that differ in which supervised checkpoint initializes reinforcement learning (RL). The largest accuracy gains from RL on compositional proofs come with the most structure beyond the atomic circuits. A circuit-level claim is therefore well defined only once one states which circuit is reported, the pruning threshold used to extract it, and the level at which circuits are compared. We distill these requirements into a reporting practice for circuit-extraction studies.

Comments26 pages, 11 figures, 20 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑