arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27158cs.LGcs.AIcs.CL

线性表示假说需要一个群作用

The Linear Representation Hypothesis Needs a Group Action

Louie Hong Yao, Yuhao Li, Shengchao Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文指出线性表示假说实为基于表示等价性的论断族,通过群作用形式化表示对象、生成过程与断言属性,并用于审计常见表示量与可解释性分析。

中文摘要 AI 辅助

要对超越特定训练模型的表示做出泛化性论断,我们需要明确两个表示何时应被视为等价。线性表示假说在讨论时往往未明确这种等价性。不同的等价概念保留不同的结构,因此,那些看似在研究同一表示的度量、探针和干预,实际上可能对应不同的假说。因此,我们认为线性表示假说并非一个单一的假说,而是一族由表示等价性区分的论断。我们利用群作用来形式化这一思想,明确表示对象、产生该对象的程序以及最终断言的属性,同时考虑模型架构所强加的等价关系。这一框架阐明了假设如何在不同度量、读取点和分析阶段之间发生变化,并利用它来审计常见的表示量以及近期可解释性分析。

英文摘要

To make claims about representations that generalize beyond a particular trained model, we need to specify when two representations should count as equivalent. The Linear Representation Hypothesis is often discussed without making this equivalence explicit. Different notions of equivalence preserve different structures, so metrics, probes, and interventions that appear to study the same representation may in fact correspond to different hypotheses. We therefore argue that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence. We formalize this idea using group actions, specifying the representation object, the procedure that produces it, and the property ultimately asserted, while accounting for equivalences imposed by the model architecture. This framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and we use it to audit common representation quantities and recent interpretability analyses.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑