arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CausalVerify:面向LLM因果推断工作流的执行落地基准

CausalVerify: End-to-End Verification of Causal Analyses by Language Models

Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie

arXiv 2609.07944首次发表:更新:

发表机构

Universidad Autónoma de Madrid; Spanish National Research Council (CSIC)(马德里自治大学; 西班牙国家研究委员会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CausalVerify通过将259篇论文与100个合成场景配对,用执行落地正确性(L2b+)评估LLM因果推断工作流,发现执行排名优于文本方向评分,且置信度不可靠。

AI 中文摘要

现有的针对大语言模型的因果推断基准大多对方法描述进行评分,或仅检查生成的代码能否运行,而不评估执行后的工作流是否恢复了目标因果估计量。CausalVerify通过将现实中的解释与可验证的计算相分离,研究结构化计量经济学因果估计工作流的这一验证问题。该基准将259篇已发表的经济学论文(包含重构的研究问题、数据描述、制度背景)与100个固定种子的合成场景配对,这些场景为双重差分、事件研究、工具变量和断点回归设计生成了CSV数据集。实验A(真实论文文本一致性)根据四个大语言模型的共识标签,对方法族和方向一致性进行评分。实验B(合成执行)运行模型编写的R代码,并检查提取的处理效应估计值是否与同一实现数据集上的规范估计量匹配;这一基于执行落地的正确性层级为L2b+,不同于仅记录代码是否执行的L2b。校准分支考察自我报告的置信度能否区分正确与不正确的工作流。在实验B中,七个大语言模型在默认50%容差下的L2b+通过率为10%至88%,在426个执行的工作流中有66个(15.5%)返回了错误的估计值。执行排名(L2b)与L2b+的一致性远优于文本方向评分(L4):Kendall τ=0.81,Spearman ρ=0.93,而L4的Kendall τ在-0.20至0.10之间。Llama-3.3-70B-Instruct表现出相同的定性差距,且报告的置信度不能可靠地区分正确与不正确的工作流。本基准的结论仅限于这四种设计族中标准化单次工作流,在评估的R后端和模型面板下;该基准不衡量通用因果推断能力。代码、数据、缓存输出和数据表均已发布。

英文摘要

Language models increasingly perform empirical analyses end to end, yet existing evaluations assess the written explanation or whether generated code executes, not whether the executed workflow recovers the intended causal estimand. We introduce CausalVerify, an execution-grounded benchmark for end-to-end causal analysis that follows a model from research-context interpretation to estimand recovery. It scores this workflow at four distinct layers: method recognition, design specification, executable implementation, and estimand recovery. It combines 259 real-paper contexts, 100 fixed-seed synthetic scenarios with executable reference estimates, and 23 paper-twin pairs in which a model commits to a design before seeing the data and its executed analysis is scored against a canonical estimator on the same realised dataset. Execution is not correctness. Among 426 model-written workflows that run without error, 15.5% fail verification, and a keyword score of the effect direction stated in the text is a poor proxy for recovery. In a 23-pair, nine-model paired study, replacing a model's committed design with the reference design and its execution conventions raises joint recovery of the point estimate and standard error from 15.0% to 51.5%; yet 48.5% of eligible seeds still fail under the reference design. The direction replicates on six pairs built afterwards under a frozen construction protocol, although on the two newest pairs the gain is confined to models from the family that built the references. Design specification is consequential but not sufficient: a plausible method and runnable code do not guarantee recovery, and even supplying the reference design and its conventions leaves substantial downstream failure. CausalVerify evaluates the executed workflow rather than its surface plausibility, and every reported number is recomputed from frozen artifacts by a single script.

Comments38 pages, 14 figures, 19 tables. v2 is a substantially revised version with a new title: it adds the layered verification framework and the paired paper-to-execution study with an oracle substitution arm. Code and data of the v1 release: https://github.com/causalverify/causalverify

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑