仅靠后期注意力层即可复制实体标记,但离不开对其上下文的关注
Late Attention Layers Alone Can Copy Entity Tokens, but Not Without Attending to Their Context
浏览论文内容
中文总结 AI 辅助
本研究通过两种新方法在Qwen3-8B上揭示,模型后半部分的两组层在上下文标记引导下对实体复制既必要又充分,且上下文标记的注意力不可或缺,确立了后期层的关键作用。
中文摘要 AI 辅助
大型语言模型(LLMs)能够可靠地执行实体复制任务,即模型将提示中指向某个实体的标记(称为实体标记)复制到其输出中以回答问题。尽管实体复制对大多数LLM而言是直接了当的,但现有研究并未系统性地说明哪些层专门负责这一基本任务,也未说明同一序列中的其他标记(称为上下文标记)如何影响模型复制实体标记的能力。为解答这些问题,我们在Qwen3-8B上开展了实验,使用了两种新方法:瓶中之灵(genie-in-a-bottle),该方法精确控制哪些层可以参与实体复制任务;以及注意力切除术(attention lobotomy),该方法在不影响其余注意力分布的前提下,切断特定标记对实体标记的注意力。我们发现,模型后半部分的两组不同层对于实体复制既是必要的也是充分的。此外,除了解码位置对实体标记的注意力外,上下文标记对实体标记的注意力对于精确复制标记也是必要的,尽管上下文标记本身并不存储实体信息,除非它们满足特定的语义属性。我们的研究结果确立了后期层在上下文标记引导下对实体复制的关键作用,并呼吁未来研究模型如何传播和消费实体信息。
英文摘要
Large language models (LLMs) reliably perform entity copying, in which a model copies tokens referring to an entity, termed entity tokens, from the prompt into its output to answer a question. Although entity copying is straightforward for most LLMs, existing research does not provide a systematic account of which layers specialize in this fundamental task or how other tokens in the same sequence, termed context tokens, influence the model's ability to copy the entity tokens. To address these questions, we conduct experiments on Qwen3-8B using two novel methods: genie-in-a-bottle, which controls exactly which layers can participate in an entity-copying task, and attention lobotomy, which cuts off specific tokens' attention to entity tokens without affecting the remaining attention distribution. We find that two distinct groups of layers in the second half of the model are both necessary and sufficient for entity copying. Moreover, in addition to the decoding position's attention to entity tokens, context tokens' attention to entity tokens also proves necessary for copying the exact tokens, even though context tokens do not store entity information themselves unless they satisfy particular semantic properties. Our findings establish the critical role of late layers in entity copying under the guidance of context tokens, calling for future work on how models propagate and consume entity information.
发表机构
- Drexel University(德雷塞尔大学)
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。