发表机构
Evident Solutions Oy(Evident Solutions 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过差分对数透镜读取Transformer内部状态,提出对比投影方法,无需训练即可追踪复合名词链,并揭示区分是跨架构稳定的,而词元表示则因网络而异。
AI 中文摘要
在词元空间中读取Transformer的内部状态容易做到,但难以信任:在中间层,对单个隐藏状态应用对数透镜时,其结果被模型对几乎任何输入都会预测的通用词元所主导。我们转而读取差异。将两个紧密匹配的提示的隐藏状态相减,并通过解嵌入进行投影,可以抵消共享成分,并凸显出将它们区分开来的内容,这一操作等同于通过对数透镜读取RepE/ActAdd引导向量。将其构建为一个无需训练的追踪器,在每一位置、子层和注意力头处进行读取,并在设计的基线上取平均,它追踪了Phi-2中一个复合名词的MLP到注意力链,并通过激活修补在该模型中得到了确认;同样的区分在三种架构中通过读出和探针而非修补得以复现;它读取了检索为真实与虚构实体所呈现的内容,并将隐喻读取为一组域到域的映射,而非单一的形象性特征。一项跨种子的对照实验划定了边界:在仅初始化不同的五个网络中,同样的区分几乎以完全不同的词元呈现(前10重叠度为0.08)。计算在词元空间中的表现是网络特定的;而它所做出的区分则不是。
英文摘要
Reading a transformer's internal states in token space is easy to do and hard to trust: a logit lens on a single hidden state is dominated, at intermediate layers, by the generic tokens the model would predict for almost any input. We read the difference instead. Subtracting two closely matched prompts' hidden states and projecting through the unembedding cancels the shared component and surfaces what separates them, an operation equivalent to reading a RepE/ActAdd steering vector through a logit lens. Built into a training-free tracer that reads at every position, sub-layer, and head and averages over designed baselines, it traces a compound- noun MLP->attention chain in Phi-2, confirmed there by activation patching, with the same distinction recovered across three architectures by readout and probe rather than by patching; it reads what retrieval surfaces for real versus fictional entities, and reads metaphor as a set of domain-to-domain mappings rather than a single figurativity feature. A cross-seed control marks the boundary: across five networks differing only in initialization, the same distinction surfaces as almost entirely different tokens (top-10 overlap 0.08). What a computation looks like in token space is network-specific; the distinction it draws is not
Comments27 pages, 34 tables. Code and data: https://github.com/EvidentSolutions/llm-interp/tree/main/contrastive