arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型内省机制的机理研究

A mechanistic study of language model introspection

Jiahong Zou, Xiangkun Sun, Lingkai Kong, Tonghan Wang

arXiv 2609.35108首次发表:更新:

发表机构

Shandong University; Tsinghua University; Northeastern University; The University of Hong Kong(山东大学; 清华大学; 东北大学; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控任务揭示LLM内省机制:中间层门控头决定是否报告内部扰动,后层路由头定位位置,且概念定位准确性影响报告精度。

AI 中文摘要

大型语言模型(LLMs)有时能够报告其内部激活的扰动——即使输入没有提供任何干预发生的证据。模型如何检测并定位此类内部变化?我们通过一项保持输入文本固定的受控任务来研究这一问题。我们要么在十个词元位置之一的隐藏状态中注入一个概念向量,要么不施加任何干预。模型被要求识别被扰动的位�置或报告未发生干预。在三个模型家族中,我们识别出两组注意力头,它们在内省报告中扮演不同角色。中间层的门控头影响模型是否报告变化,而较后层的路由头帮助选择要报告的位置。对门控头的干预可以抑制位置报告,即使路由头提供了位置信息。我们进一步考察了为何报告准确性在不同概念间存在差异。定位更准确的概念向量在门控头中产生更强的注意力分数和输出响应,这与它们QK和OV计算中诱导的键和值变化更好地对齐有关。总之,这些发现揭示了支持内省检测和定位的注意力头机制。

英文摘要

Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑