发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体在上下文变化时使用旧值回答的“陈旧绑定”问题,提出CICM基准并揭示注意力漂移机制,通过重定向注意力无需训练即可纠正多数错误。
AI 中文摘要
随着偏好、目标和事实的变化,大语言模型(LLM)智能体必须使用当前状态,而早期版本仍保留在上下文中。然而,它们可能会用同一变量的旧值来回答,我们将这种失败称为“陈旧绑定”(stale binding)。为了研究模型何时使用过时信息以及为何如此,我们引入了受控上下文记忆(Controlled In-Context Memory, CICM),这是一个用于在对话和智能体日志中跟踪和使用更新信息的基准。我们观察到,即使是前沿推理模型也可能无法恢复当前状态。我们发现,在开源模型中,当模型用旧值回答时,探针(probes)仍然能够恢复更新后的值,这表明模型未能选择仍然可用的信息。对Qwen和Pythia的组件测试识别出这种选择失败的一种机制:注意力漂移(attention drift),即在生成答案时,注意力偏向旧值而非当前值。我们研究了一个单层Transformer,从数学上理解这一现象是如何发生的:当注意力分数相似时,多个旧值可以共同获得比当前值更多的注意力。在这一解释的指导下,我们在不进行额外训练的情况下,将注意力重新导向当前值。当直接请求当前值时,针对每个输入调整这种干预,可以在各种模型家族中纠正大多数旧值错误,同时几乎保留所有最初正确的答案。因此,可靠的上下文管理不仅仅需要记住更新的信息:模型必须使用它来指导其答案。
英文摘要
As preferences, goals, and facts change, LLM agents must use the current state while earlier versions remain in context. Yet they can answer with an old value of the same variable, a failure that we call stale binding. To study when models use outdated information and why, we introduce Controlled In-Context Memory (CICM), a benchmark for tracking and using updated information in conversations and agent logs. We observe that even frontier reasoning models can fail to recover the current state. We find that in open-source models probes can still recover the updated value when the model answers with an old one, pointing to a failure to select information that remains available. Component tests in Qwen and Pythia identify a mechanism for this selection failure: attention drift, where attention favors old values over the current one when producing an answer. We study a one-layer transformer to mathematically understand how this phenomenon happens: when attention scores are similar, several old values can together receive more attention than the current value. Guided by this explanation, we redirect attention toward the current value without further training. When the current value is requested directly, adjusting this intervention for each input corrects most old-value errors across various model families while preserving nearly all initially correct answers. Reliable context management therefore requires more than remembering updated information: models must use it to guide their answers.