arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31675cs.LGcs.AI

主动因果发现基准:评估预算干预下的LLM智能体

Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions

Sagar Deb, Devam Shah, Ashwanth Krishnan

首次发表
浏览论文内容

中文总结 AI 辅助

提出ACDB基准,评估LLM智能体在预算干预下恢复因果图的能力,发现PC算法优于LLM,且当前LLM未解决主动因果发现,结果作为审计与校准报告。

中文摘要 AI 辅助

我们引入了主动因果发现基准(ACDB),这是一个基于结构方程模型(SCM)的环境,用于评估LLM智能体能否从观测数据和预算受限的硬干预中恢复因果图结构。ACDB将线性高斯世界生成器与固定的观测-干预-提交API以及三层评分合约相结合,该合约分别评估骨架恢复、DAG恢复和干预效率。在当前六级阶梯上,采用贪婪主动定向启发式的PC算法是最强的非预言机方法(定向F1为42.7%,SHD为4.79),领先于Claude Sonnet 4.6原始主动方法(31.7%,7.25)和GPT-5.4原始主动方法(22.9%,9.27)。最具信息量的诊断是精确率-召回率分解:PC算法高精确率但干预不足,LLM则过度干预且精确率较低,而统计工具访问往往增加弃权(不执行)而非有用的干预。在此密集v0阶梯上,结构盲随机DAG基线达到23.6%的定向F1;密度探针将该下限降至16.9%,促使v1校准通过。因此,当前结果应被视为基准审计和校准报告,而非表明现有LLM已解决主动因果发现问题的证据。

英文摘要

We introduce the Active Causal Discovery Benchmark (ACDB), an SCM-grounded environment for evaluating whether LLM agents recover causal graph structure from observations and budget-constrained hard interventions. ACDB pairs a linear-Gaussian world generator with a fixed observe-intervene-submit API and a three-layer scoring contract that separates skeleton recovery, DAG recovery, and intervention efficiency. On the current six-level ladder, PC with a greedy active orientation heuristic is the strongest non-oracle method (directed F1 42.7%, SHD 4.79), ahead of Claude Sonnet 4.6 raw active (31.7%, 7.25) and GPT-5.4 raw active (22.9%, 9.27). The most informative diagnostic is the precision-recall decomposition: PC under-commits with high precision, LLMs over-commit with lower precision, and statistical-tool access often increases abstention rather than useful intervention. A structure-blind random DAG baseline reaches 23.6% directed F1 on this dense v0 ladder; a density probe lowers this floor to 16.9%, motivating the v1 calibration pass. The current results should therefore be read as a benchmark audit and calibration report, not as evidence that current LLMs solve active causal discovery.

↑