arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39263cs.CLcs.LG

概念子空间超越 Logit Lens 的计算:一种仅基于权重的测试,用于定位读出上游的表示

Concept Subspaces Compute Beyond the Logit Lens: A Weights-Only Test for Locating Representations Upstream of Readout

Aojie Yuan, Zhiyuan Julian Su, Haiyue Zhang, Zijian Su

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一种仅基于权重的几何诊断方法,通过测量概念子空间与解嵌入矩阵主奇异方向的重叠,区分概念结构与读出方向,并验证了提取过程的迁移性,限制了因果充分性主张。

中文摘要 AI 辅助

概念子空间对模型行为的影响并不能确定其与输出读出的关系。我们引入了一种双面几何诊断方法,该方法测量提取的子空间与解嵌入矩阵的 dominant 右奇异方向的重叠,并针对输出导向的阳性对照进行评估。给定提取的基,原始诊断仅需要模型权重。我们的测试平台是格式无关推理子空间(FARS),这是一个从以六种表面形式表达的十八个推理概念中提取的十维基。在九个秩匹配估计器和二十六个模型中,四个激活导出的概念估计器在 top-ten 读出跨度中仅携带 0.38--0.80% 的平均能量。最终层 PCA 携带 3.56%,在 26 个模型中的 25 个中超过 FARS。使用拟合的线性翻译器进行深度匹配评估的同一层 next-token 对照,携带的能量约为 FARS 的十三倍,在所有 25 个测试模型中均有分离。在十个不相交的概念上重新提取 FARS,在二十四个生成模型中产生 62--100% 的跨格式检索,证明了提取过程的迁移而非固定基。一项互补的四模型、三种子干预研究发现模型依赖的源定向效应,这些效应仍远低于全向量替换。总之,几何和干预对照将概念结构与 dominant 读出方向区分开来,同时限制了因果充分性的主张。

英文摘要

A concept subspace's effect on model behavior does not establish how it relates to the output readout. We introduce a two-sided geometric diagnostic that measures an extracted subspace's overlap with the dominant right-singular directions of the unembedding matrix, evaluated against output-oriented positive controls. Given an extracted basis, the raw diagnostic requires only model weights. Our testbed is the Format-Agnostic Reasoning Subspace (FARS), a ten-dimensional basis extracted from eighteen reasoning concepts expressed in six surface forms. Across nine rank-matched estimators and twenty-six models, four activation-derived concept estimators carry only 0.38--0.80% mean energy in the top-ten readout span. Final-layer PCA carries 3.56%, exceeding FARS in 25 of 26 models. A same-layer next-token control, evaluated using a fitted linear translator for depth matching, carries approximately thirteen times more energy than FARS, with separation in all 25 tested models. Re-extracting FARS on ten disjoint concepts yields 62--100% cross-format retrieval across twenty-four generative models, demonstrating transfer of the extraction procedure rather than a fixed basis. A complementary four-model, three-seed intervention study finds model-dependent source-directed effects that remain well below full-vector replacement. Together, the geometry and intervention controls distinguish concept structure from dominant readout directions while limiting claims of causal sufficiency.

发表机构

  • University of Southern California(南加州大学)
  • Duke University(杜克大学)
  • University of Michigan(密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑