发表机构
Saarland University; Hamburg University of Technology; Helmholtz-Zentrum Hereon; German Research Centre for Artificial Intelligence (DFKI)(萨尔大学; 汉堡工业大学; 亥姆霍兹中心赫伦; 德国人工智能研究中心(DFKI))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文推出针对大语言模型锚定效应的基准测试AnchorBench,经14种模型实验,揭示锚定效应的路径依赖性等特性,发现高控制准确率的前沿模型仍易受合理锚点影响。
AI 中文摘要
锚定效应是一种认知偏差,指初始参考值会将后续判断向自身偏移。该效应在人类判断与决策中已得到充分证实,近期研究显示大语言模型(LLM)也存在类似行为。然而,现有针对LLM锚定效应的研究通常仅评估有限的锚定路径,且极少区分无关锚点与合理锚点。本文推出AnchorBench,这是一个针对LLM锚定效应的基准测试,在明确的锚点相关性维度下评估多种锚定路径。研究涵盖14种模型,包括10种开源权重模型和4种前沿API模型,以及大量受控提示,发现:(1)锚定效应具有强烈的路径依赖性;(2)当合理锚点通过更强路径引入时,通常会比无关锚点引发更大的偏移;(3)随着锚点与证据支持的答案距离越远,锚点影响通常会减弱,在外部锚点和检索增强生成(RAG)上表现最为明显;(4)无锚点控制条件下的高任务准确率(Acc₁₀:与标准答案相差10分以内的答案)无法保证鲁棒性:即使控制准确率超过95%的前沿API模型,仍易受合理锚点影响。
英文摘要
The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language models (LLMs) exhibit similar behavior. However, existing work on anchoring in LLMs typically evaluates only a narrow set of anchor pathways and rarely distinguishes irrelevant from plausible anchors. We introduce AnchorBench, a benchmark for the anchoring effect in LLMs that evaluates multiple anchor pathways under an explicit anchor relevance axis. Across fourteen models, including ten open-weight models and four frontier API models, and a large set of controlled prompts, we find that (1) anchoring is strongly pathway-dependent, (2) plausible anchors usually induce larger shifts than irrelevant ones when introduced through stronger pathways, (3) anchor influence generally weakens as the anchor moves farther from the evidence-supported answer, most clearly on External and RAG, and (4) high task accuracy on the anchor-free control condition (Acc$_{10}$: answers within 10 points of gold) does not guarantee robustness: even frontier API models above 95% control accuracy remain susceptible to plausible anchors.
CommentsPublished as a conference paper at COLM 2026