发表机构
Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机制可解释性缺乏有界邻域保证的问题,提出基于约束多项式-区域传播的框架,将单输入发现提升为认证声明,并实例化于变压器注意力以锐化机制结论。
AI 中文摘要
机制可解释性每次针对一个输入对变压器电路进行逆向工程,导致所观察到的机制在有界输入邻域上缺乏保证。我们通过一个基于约束多项式-区域(CPZ)传播的框架来解决这一差距,该框架将机制可解释性观察从单个输入提升为对有界扰动集的认证声明。三个内部注意力查询(top-k稳定性、证据质量和注意力熵)被表述为注意力权重单纯形上的可处理程序,并且通过变压器块的CPZ传播被证明能精确保持softmax单纯形和LayerNorm零均值恒等式。递归雅可比区域构造通过在输入处对块堆栈进行线性化,将相同的认证扩展到层深度,并避免逐层生成器增长。我们在变压器注意力上实例化该框架;由此产生的认证提供了一种方式来锐化单输入检查无法自行解决的机制声明,并在经验启发式可能误导的情况下为下游决策提供信息。
英文摘要
Mechanistic interpretability reverse-engineers transformer circuits one input at a time, leaving observed mechanisms without guarantees over bounded input neighbourhoods. We address this gap with a framework based on constrained polynomial-zonotope (CPZ) propagation that lifts mechanistic-interpretability observations from a single input to certified statements over a bounded set of perturbations. Three internal-attention queries (top-$k$ stability, evidence mass, and attention entropy) are formulated as tractable programs over the simplex of attention weights, and CPZ propagation through transformer blocks is shown to preserve the softmax simplex and the LayerNorm zero-mean identity exactly. A recursive Jacobian zonotope construction extends the same certificates across layer depth by linearising the block stack at the input and avoids per-layer generator growth. We instantiate the framework on transformer attention; the resulting certificates offer a way to sharpen mechanistic statements that single-input inspection cannot resolve on its own, and to inform downstream decisions in regimes where empirical heuristics may be misleading.