arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

答案盆地表示假说:我们并非在探测或引导概念

The Answer-Basin Representation Hypothesis: We Are Not Probing or Steering Concepts

Manjiang Yu, Hongji Li, Zihan Wang, Junwei Chen, Xue Li, Priyanka Singh, Yang Cao, Lijie Hu

arXiv 2609.24821首次发表:更新:

发表机构

The University of Queensland; Mohamed bin Zayed University of Artificial Intelligence; Institute of Science Tokyo(昆士兰大学; 穆罕默德·本·扎耶德人工智能大学; 东京科学大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出答案盆地表示假说,认为语言模型中概念相关线性结构源于答案概率测度差异,实验验证了探测和引导效应与概念标签和答案测度的对齐。

AI 中文摘要

线性表示假说将高层概念与语言模型中的方向相关联,但这些与概念相关的线性结构在模型内部如何组织仍不清楚。我们提出答案盆地表示假说:模型续写分布诱导的答案上的概率测度组织这些线性结构,其统计量沿跨问题共享的线性方向表示。产生相同答案的所有续写构成一个答案盆地,其质量即它们的总概率。这些盆地质量定义了答案上的前推概率测度。我们认为,与概念相关的线性结构源于答案测度的差异,而非由概念标签的变化决定。跨模型和任务的实验将探测和引导中概念一致效应及其反转与概念标签和答案测度之间的对齐联系起来。

英文摘要

Linear probing and activation steering use linear directions to predict and control behavior-level concepts such as correctness, safety, and social bias in question answering. We call these \emph{behavior-level concept directions}. Drawing on the Linear Representation Hypothesis (LRH) and circuit studies, these directions are often interpreted as concept representations. Yet controlled evidence for linear concept structure and local mechanisms does not establish that these directions represent the intended behavioral concepts, leaving probing and steering without a unified theoretical account. We propose the \emph{Answer-Basin Representation Hypothesis} (ABRH): the model's own answer measure organizes the linear structure of these directions. For each question, all continuations yielding the same answer form an answer basin, whose mass is their total probability; these masses define the model's answer distribution. ABRH posits that its concentration before generation and the relative mass of each answer after generation are represented along linear directions shared across questions. Experiments span four Qwen2.5 and Gemma-3 models on correctness, social bias, and safety tasks. Decoupling concept labels from basin mass shows that probing and steering exhibit concept-consistent effects when labels align with mass orderings, weaken as mass gaps shrink, and reverse under conflict. Probes trained solely on mass orderings among wrong answers still select correct answers; a fixed steering vector can instead favor wrong answers when labels conflict with mass orderings. These results support an answer-measure account of behavior-level concept directions and of when probing and steering succeed, fail, or reverse.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑