当模型不操纵流形时:比较任务的几何结构
When Models Don't Manipulate Manifolds: The Geometry of a Comparison Task
浏览论文内容
中文总结 AI 辅助
本研究通过分析Qwen2.5-7B-Instruct在数字比较任务中的因果几何结构,发现模型主要依赖线性表示而非流形结构进行计算,表明流形假设与线性表示可共存,为机制可解释性提供了新见解。
中文摘要 AI 辅助
机制可解释性研究当前的一个前提是,对神经网络表示几何结构的详细描述可以告诉我们模型如何执行计算,以及如何有效地对它们进行干预。尽管文献中已观察到多个概念的低维流形(例如,数字编码在螺旋上,星期几编码在圆上……),且其结构被认为反映了数据和任务的性质,但模型在多大程度上依赖这些流形进行计算,以及它们如何操纵这些流形,仍不清楚。我们精确刻画了数字比较任务中计算的几何结构,将其作为决策中比较的一种抽象,以及模型如何以优雅的方式利用几何结构来实现该任务。具体而言,我们研究了Qwen2.5-7B-Instruct(一个能力强且被广泛研究的开源权重模型)中数字比较的因果几何结构,发现Qwen在很大程度上使用数字的线性表示,尽管存在弯曲的几何结构。为了比较两个数字,模型首先将每个数字编码为一个向量,并通过注意力和残差连接将两个表示相加,将它们带入残差流中的共享空间。然后,模型使用MLP神经元在该共享空间的局部区域(对应于输入数字的较小区间)比较这对数字,并组合这些结果以获得最大值的位置。事实上,这种对线性表示进行比较的依赖在模型比较三个数字时也持续存在。我们的发现表明,流形假设可以与线性表示共存:虽然有序概念在表示中可能具有流形结构,但模型在某些计算中可能使用该概念的底层线性结构。
英文摘要
One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can tell us how models perform computations, and how to effectively intervene on them. While low dimensional manifolds have been observed for multiple concepts in the literature (e.g. numbers encoded on helices, days of the week on a circle, ...), with structure believed to reflect properties of data and tasks, the extent to which models rely on them for computation, and how they manipulate them, remains unclear. We characterize precisely the geometry of computation in a number-comparison task, as an abstraction of comparison for decision making, and how models utilize geometry in an elegant fashion to implement it. Specifically, we study the causal geometry of number comparison in Qwen2.5-7B-Instruct, a capable and widely studied open-weight model, and find Qwen largely uses linear representations of numbers despite the presence of curved geometry. To compare two numbers, the model first encodes each number along a vector and adds the two representations using attention and the residual connection, bringing them into a shared space in the residual stream. Then, the model uses MLP neurons to compare the pair of numbers on local regions in this shared space, which correspond to smaller intervals of input numbers, and combines these to obtain the position of the maximum. In fact, this reliance on linear representations for comparison also persists when the model compares three numbers. Our findings demonstrate that the manifold hypothesis can co-exist with linear representations: while concepts that are ordered may have manifold structure in representations, the model may use an underlying linear structure of the concept in certain computations.
发表机构
- Goodfire AI
机构由 AI 辅助整理,请以论文原文为准。