arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33065cs.CLcs.LG

过度解读上下文:被动暴露可左右大语言模型决策

Reading Too Much into Context: Passive Exposure Can Steer LLM Decisions

Yuxiang Zheng, Lin Tian, Marian-Andrei Rizoiu

AI总结:

研究发现,即使外部内容不提供改变理由,被动暴露于额外上下文也能系统性左右大语言模型的决策,在闭源模型中影响可达近50个百分点,甚至导致违反用户要求或接受虚假声明。

AI中文摘要:

大语言模型(LLM)助手在完成用户请求时,现在可以搜索网络并查阅外部来源。这些来源可以提供有用的证据,但也可能向模型的上下文引入额外内容。这种被动暴露是否能在新增内容不提供任何改变理由的情况下左右决策?我们考察了在有无此类外部内容的情况下,模型在同一任务上决策的稳定性。在我们测试的所有开源权重和闭源权重模型中,暴露系统性改变了决策,在闭源权重模型中影响幅度接近50个百分点。同样的模式也出现在真实世界的在线观点中。这种影响还超出了主观偏好。此类暴露可以引导模型选择违反明确用户要求的选项,并提高其对虚假声明的接受度。简而言之,进入LLM上下文的内容即使不应决定决策,也能影响其决策。

英文摘要:

Large language model (LLM) assistants can now search the web and consult external sources while completing user requests. These sources can provide useful evidence, but they can also introduce additional content into the model's context. Can such passive exposure steer a decision even when the added content provides no reason to change it? We examine the stability of model decisions on the same tasks with and without such external content. Across all open-weight and closed-weight models we test, exposure systematically shifts decisions, with effects reaching nearly 50 percentage points in closed-weight models. The same pattern appears with real-world online opinions. The influence also extends beyond subjective preferences. Such exposure can steer models toward choices that violate explicit user requirements and increase their acceptance of false claims. In short, what enters an LLM's context can influence its decision even when it should not determine it.

补充信息

↑