arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00753cs.LGcs.CL

语言模型如何在上下文与记忆之间做出选择?

How Do Language Models Choose Between Context and Memory?

发表机构斯坦福大学 · 永久实验室
查看机构详情
  • Stanford University(斯坦福大学)
  • Perpetual Labs(永久实验室)

机构由 AI 辅助整理,请以论文原文为准。

Benjamin Shih, John Winnicki, Arianna Cao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过反事实实验,在Qwen、Llama、OLMo模型上揭示语言模型的上下文与参数知识选择受任务依赖的权威方向影响,局部方向对源选择转变的重现效果优于跨任务方向。

中文摘要 AI 辅助

当上下文信息与模型参数中存储的知识发生冲突时,激活方向可用于解码并引导模型选择遵循哪一来源。然而,沿某一方向进行引导并不能确立因果关系:未经过编辑的模型是否会自然使用该方向,或者该方向是否可跨任务复用。我们在明确的设置中通过反事实实验检验这些区别。首先,我们从一致性提示中估计权威方向,其中上下文与参数化知识支持相同答案。随后,我们在匹配的提示之间交换沿这些方向自然出现的坐标,这些提示会引导模型优先选择所提供的上下文或其参数化知识。在Qwen、Llama和OLMo模型上,该干预措施可重现30%-68%的权威诱导源选择转变,而匹配的对照组几乎无法重现任何转变。为测试跨任务复用,我们分别在两个任务上学习权威方向,发现跨任务可迁移性仅缩小了9%的权威差距,而在给定任务上学习的局部方向缩小了57%的差距。这些结果区分了权威表示、因果使用和跨任务因果复用,表明权威计算可能依赖于任务,而非可跨任务复用。

英文摘要

When contextual information conflicts with knowledge stored in model parameters, activation directions can be used to decode and steer which source the model follows. However, successful steering does not establish that the unedited model uses those directions to choose between sources, or that they remain effective across tasks. To test these possibilities, we vary the stated authority of contextual claims while holding their content fixed. We first estimate authority directions from prompts in which context and parametric knowledge agree, then test their causal contribution when the two sources conflict. Interchanging naturally occurring activation values along these directions between matched high- and low-authority prompts reproduces 30--68% of the authority-induced shift in source choice across Qwen, Llama, and OLMo models, whereas matched controls reproduce almost none. We next ask what transfers across tasks: the learned direction versus the activation values exchanged along it. Using a direction learned on another task closed 9% of the source-choice gap, compared with 57% when learned on the task being evaluated. Both interventions exchanged activation values from the evaluated task. In a separate experiment, we kept its learned direction but exchanged values taken from another task, which closed 68% of the gap. These results show that authority-related activation values can causally influence source choice across tasks when inserted along directions learned for the task being evaluated.

↑