你会步行去洗车吗?揭示大型语言模型在常识推理中的显著性偏差
Would You Walk to the Car Wash? Salience Bias in LLM Commonsense Reasoning
AI总结:
本研究构建SaliTrap基准数据集,发现主流LLMs存在显著性偏差,其常识推理失败源于知识被干扰项抑制,而非知识缺失,轻量级推理时提示可大幅缓解该问题。
AI中文摘要:
随着大型语言模型(LLMs)在复杂推理任务中不断进步,它们已学会高度优先考虑输入中提供的显式条件。然而,在日常常识推理中,这种机制暴露出一个关键漏洞,我们将其称为显著性偏差:模型极易被无用的显式干扰项(例如数值)误导,从而忽略任务隐含的物理或常识前提。一个关键的开放性问题是,这种失败是否反映了常识知识的真正缺口,还是仅仅是在误导性任务框架下的知识被抑制。为了探究这一点,我们构建了SaliTrap基准(SaliTrap Benchmark),这是一个涵盖四个陷阱维度的高质量数据集。对12个最先进的LLMs进行评估,我们发现所有主流模型都存在显著的显著性偏差,其严重程度随干扰项密度增加而上升,且检测到陷阱与实际避开陷阱往往相互脱节。至关重要的是,通过剥离任务框架后重新激发相同模型,我们表明这主要是知识抑制而非知识缺失的失败:仅无上下文的知识探测就能恢复超过90%的谄媚式顺从失败,表明必要的常识本质上存在,但被显著性干扰项主动排挤,这些干扰项诱使模型采取过度顺从、不必要的计算。基于这一诊断,我们进一步证明,仅轻量级的推理时提示就能大幅缩小差距,无需任何再训练。我们的发现将常识推理失败的瓶颈从模型能力转移到激发方式,并发布SaliTrap作为这一盲点的测试平台。代码可在该https URL获取。
英文摘要:
Despite advances in complex reasoning, large language models (LLMs) can prioritize explicit input conditions over implicit task prerequisites. In everyday commonsense reasoning, this can lead to a failure we term Salience Bias: salient but task-irrelevant details (e.g., numerical values) draw models into computation while they overlook the physical or commonsense prerequisites of the task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, comprising 1,145 items in four trap dimensions. Evaluating 12 LLMs, we find substantial vulnerability across the tested models, with higher numerical distractor counts associated with lower trap-avoidance rates and trap recognition not always leading to avoidance. Further probing of sycophantic-compliance cases shows that the relevant commonsense can often be elicited when the original task framing is removed, suggesting a gap between recognizing a constraint and applying it during task execution. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings highlight the importance of applying commonsense constraints during task execution, and we release SaliTrap as a testbed for studying this gap. The codes are available at https://github.com/Wuzheng02/SaliTrap