发表机构
University of Utah; Bangladesh University of Engineering and Technology(犹他大学; 孟加拉工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出DeReLab生成式基准框架,评估9种大型语言模型,发现其普遍存在确认偏差,即易接受一致证据、抵制不一致更新,部分模型能识别弱化更新却不修正结论。
AI 中文摘要
可废止推理是一种基于当前合理证据得出推论,但在引入新证据时可撤回推论的推理类型。尽管近期研究已探究语言模型在可废止推理中的行为,但所用数据集为静态的,且缺乏对非单调推理类别的广泛覆盖。我们推出DeReLab,这是一个生成式框架,可从默认和继承推理的参数化图结构中生成多轮信念更新对话,每轮均有经形式验证的真实值,从而能受控地测量模型对确认性和否定性证据的响应。这种受控生成过程为隔离特定推理需求的实验设计创建了测试平台。将此能力应用于确认偏差研究,我们评估了9种开源和专有大型语言模型,发现几乎所有模型都表现出系统性倾向:接受一致证据,同时抵制不一致更新,部分模型能正确识别弱化更新,但未能修正其结论。我们认为,本工作及发现将推动未来对语言模型可废止推理评估的研究。
英文摘要
Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have examined language-model behaviors in defeasible reasoning, the datasets have been static and lack wide coverage of non-monotonic reasoning categories. We introduce DeReLab, a generative framework that produces multi-turn belief-updating conversations from parameterized graph structures across default and inheritance reasoning, with formally verified ground truth at every turn, enabling controlled measurement of how models respond to confirming and disconfirming evidence. This controlled generation process creates a testbed for experimental designs that isolate specific reasoning demands. Applying this capability to the study of confirmation bias, we evaluate nine open and proprietary large language models and find that nearly all exhibit a systematic tendency to accept congruent evidence while resisting incongruent updates, with several models correctly identifying a weakening update yet failing to revise their conclusion. We believe our work and findings will facilitate future research on evaluating language models in defeasible reasoning.
CommentsAccepted at the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)