arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27953cs.AI

《“如果”的错觉:评估大型语言模型中反事实推理的失效》

The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs

  • Zhejiang University(浙江大学)
  • Alibaba Group(阿里巴巴集团)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

Yucheng Wang, Yuetian Du, Zhengyi Liu, Rongyu Zhang, Bing Zhao, Boyu Yang, Ming Kong, Lin Qu, Hu Wei, Jie Liu, Qiang Zhu

AI总结:

本研究提出开放域反事实推理基准WhatIfBench及评估框架PRISM,发现前沿LLM在该基准上表现未饱和,存在因果推理缺陷,凸显其反事实推理能力的不足。

AI中文摘要:

反事实推理要求模型超越观测到的世界进行推理,并解释改变的条件如何传递到下游结果。现有基准大多针对变量固定或仅有单一黄金结果的受限场景,忽略了需要因果过程评估的开放域场景。为此,我们提出了**WhatIfBench**,一个针对开放域、开放形式、长跨度反事实因果推理的诊断基准,包含STEM、HSS和混合场景下的220个“如果”问题。为评估自由形式的回答,我们进一步提出了**PRISM**,它首先将每个自然语言解释转换为事件、状态和机制的响应衍生语义因果图。在此图基础上,PRISM随后联合应用评估图级因果有效性的过程指标,以及评估回答级解释充分性的准则指标。使用该框架评估六个前沿大型语言模型后,我们发现WhatIfBench远未达到饱和:即使是最强的模型也仅达到64.62%的最终分数。进一步分析揭示了持续存在的因果缺口、前提漂移和拓扑碎片化,表明流畅的反事实叙事往往掩盖了脆弱的因果过程。该基准、代码和评估脚本可在$\textbf{WhatIfBench}$获取。

英文摘要:

Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks largely target bounded settings with fixed variables or single gold outcomes, overlooking open-domain scenarios requiring causal-process evaluation. To this end, we present $\textbf{WhatIfBench}$, a diagnostic benchmark for open-domain, open-form, long-horizon counterfactual causal reasoning, containing 220 what-if questions across STEM, HSS, and Hybrid scenarios. To evaluate free-form responses, we further propose $\textbf{PRISM}$, which first converts each natural-language explanation into a Response-Derived Semantic Causal Graph of events, states, and mechanisms. On top of this graph, PRISM then jointly applies a Process Metric assessing graph-level causal validity and a Rubric Metric assessing answer-level explanatory adequacy. Evaluating six frontier LLMs with this framework, we find that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score. Further analysis reveals persistent causal gaps, premise drift, and topology fragmentation, suggesting that fluent counterfactual narratives often mask fragile causal processes. The benchmark, code, and evaluation scripts are available at $\href{https://github.com/zju-gt/WhatIfBench}{WhatIfBench}$.

补充信息

↑