arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

编码早,使用晚:Transformer 在何处开始对推断出的合作伙伴专业知识采取行动

Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise

Mika Okamoto, Gabriele Sarti

arXiv 2609.07139首次发表:更新:

发表机构

Georgia Institute of Technology; Northeastern University(佐治亚理工学院; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过ExpertCollab语料库发现,Transformer对对话伙伴专业水平的推断在早期层即可解码,但直到网络后期才因果影响输出,为干预时机提供了界限。

AI 中文摘要

Transformer 可以在其残差流中,在某个属性尚未影响输出的深度上,使该属性线性可解码。这种信息可读位置与使用位置之间的差距,已被证明适用于输入中直接陈述的属性。我们探究这一现象是否也适用于模型必须在对话中逐步推断的属性,即其对话伙伴的专业水平。利用 ExpertCollab,一个包含四个专业水平的模型扮演角色之间多轮研究规划对话的语料库,我们发现伙伴的专业水平在早期层中最易解码,并在网络中点之前降至接近随机水平。反事实修补表明,在峰值可解码层注入专业水平差异几乎不会改变固定的后期层读出,而在中点之后注入相同差异则几乎完全传播,两者相差超过一个数量级。内容匹配的随机对照和无探针诊断将转变点定位在同一早期层,而静态指定的对照属性则在整个网络中保持可解码。因此,推断出的关系属性在其因果激活之前就已得到充分表征,这限制了任何试图读出或引导伙伴条件行为的干预必须发生的位置。我们使用一个模型在合成语料库上进行初步演示。

英文摘要

A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This gap between where information is readable and where it is used has been shown for attributes stated directly in the input. We ask whether it also holds for an attribute the model must infer gradually over a conversation, namely how expert its dialogue partner is. Using ExpertCollab, a corpus of multi-turn research-planning dialogues between model-played personas at four expertise levels, we find that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network. Counterfactual patching shows that injecting the expertise difference at the layer of peak decodability barely changes a fixed late-layer readout, whereas the same difference injected past the midpoint propagates almost completely, a separation of more than an order of magnitude. A content-matched random control and a probe-free diagnostic place the transition at the same early layer, and a statically specified control attribute stays decodable throughout. An inferred relational attribute is therefore represented well before it becomes causally active, which bounds where any attempt to read out or steer partner-conditioned behavior must intervene. We use one model on a synthetic corpus as an initial demonstration.

CommentsPublished at the Scientific Understanding of Foundation Models (Sci-FM) Workshop at COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑