聚集而非准入:注意力如何将隐变量转化为可表述形式
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
浏览论文内容
中文总结 AI 辅助
该研究通过实验发现语言模型不存在词工作空间理论预测的准入门,揭示了隐变量的可读形式由中层窗口内注意力介导的需求特异性聚集产生,而非门控机制。
中文摘要 AI 辅助
语言模型以可报告的形式存储隐式量,当任务需要灵活复用该量时,该形式中会呈现更多相关量。表征进入该形式的原因尚不明确,“词工作空间”理论提出了“准入门”机制,即存在一个门控单元决定哪些内容可进入该形式。本研究使用雅可比透镜(Jacobian lenses)对开放权重模型进行测试,该模型的五个分支共享完全相同的上下文。研究发现,不存在符合预测的门控单元。在主检查点上,需求使某概念的透镜可见度超出对提供值应用算子所产生的结果,百分位排名提升+0.050(95%置信区间[+0.045, +0.057]),在测量的四个检查点上均为正值,尽管该分支的回答已达天花板水平,且在该读出方式下,与准确率匹配的对照效应更强。同时,存在一个共享的线性映射从包括对照分支在内的所有分支中解码该变量,其强度为经选择校正的基线的6.4至9.0倍。在查询位置产生后期可读形式的原因是中层窗口内的注意力介导聚集:将 patch 深度与读出深度分离后,在非饱和读出条件下,该窗口内的传输强度至少比任何更浅位置高17倍,且测试的 MLP 输出在该窗口内无正向贡献。在饱和百分位排名下,同一网格无法定位该窗口,这是该测量指标的特性。无需该变量的分支的聚集强度低7倍,因此该窗口具有需求特异性。该窗口有两个测量边界,下方为存活失效,上方为破坏,且在64层混合模型与另一系列的62层密集模型中,该窗口位于相同的分数深度处。本研究定位了变量的存储与读取位置,而非传输路径,该路径不传输任何内容。但读出并非使用的校准测量:三个组件使读出结果彼此相差不超过12%,但它们对答案的影响差异达7.4倍。
英文摘要
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open-weight models with Jacobian lenses, over a benchmark whose five arms share an identical context, we find no gate where it predicts one. Demand raises a concept's lens visibility beyond what applying an operator to a supplied value produces: +0.050 [+0.045, +0.057] in percentile rank on our primary checkpoint, positive on all four we measure, though that arm answers at ceiling and the accuracymatched contrast is stronger under that readout. At the same time one shared linear map decodes the variable from every arm, the control included, at 6.4-9.0x its selection-corrected floor. What produces the later readable form at the queried position is attention-mediated gathering inside a mid-depth window: separating patch depth from readout depth puts transport there at least 17x above anywhere shallower under non-saturating readouts, with no tested MLP output contributing positively inside it. Under the saturating percentile rank the same grid does not localise the window, which is a fact about that measure. An arm that needs the variable for nothing concentrates sevenfold less, so the window is demand-specific. That window has two measured edges, a survival failure below and destruction above, and it falls at the same fractional depth in a 64-layer hybrid and a 62-layer dense model from another family. We localise where the variable is installed and read, not the route from the passage, which transports nothing. But the readout is not a calibrated measure of use: three components move it to within 12% of one another and differ 7.4x in what they do to the answer.
发表机构
- University of California, Santa Cruz(加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。