CausalArena:基础模型时代的因果发现基准测试
CausalArena: Benchmarking Causal Discovery in the Foundation Model Era
- School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
CausalArena提出统一可演进的基准,通过合成、语义操作和公式基础SCMs及真实数据评估因果发现,揭示跨基准性能不可迁移,强调多样性与预训练重叠挑战。
AI中文摘要:
因果发现旨在从数据中揭示因果结构,是科学推理和基于干预的决策的基础。其评估在很大程度上依赖于结构因果模型(SCMs),这些模型指定了因果图以及生成数据的机制,然而现有研究在图族、机制和评估协议方面存在显著差异。因果发现基础模型(CDFMs)的出现进一步使评估复杂化:性能可能不仅反映因果发现能力,还反映预训练环境与测试SCMs之间的重叠,这使得在固定合成基准上的结果难以解释。我们引入了CausalArena,一个在统一协议下进行因果发现的统一且可演进的基准。合成SCMs在结构和机制上提供了受控的广度;语义操作型SCMs提供了超越标准合成生成器的人类可审计、语义基础的环境;而公式基础型SCMs则在明确的科学机制下测试发现能力。公开的真实世界数据集提供了额外的外部有效性检查。在经典方法、神经方法和预训练方法上的实验揭示了在不同SCM族和协议间的显著排名变化,表明在一个基准场景中的强性能并不能可靠地迁移到其他场景。这些结果凸显了基准多样性和预训练-评估重叠作为基础模型时代评估因果发现的核心挑战。
英文摘要:
Causal discovery aims to uncover causal structures from data and is fundamental to scientific reasoning and intervention-based decision making. Its evaluation relies heavily on structural causal models (SCMs), which specify a causal graph together with the mechanisms that generate data, yet existing studies differ substantially in graph families, mechanisms, and evaluation protocols. The emergence of causal discovery foundation models (CDFMs) further complicates evaluation: performance may reflect not only causal discovery ability, but also overlap between pretraining environments and test SCMs, making results on fixed synthetic benchmarks difficult to interpret. We introduce CausalArena, a unified and evolvable benchmark for causal discovery under a common protocol. Synthetic SCMs supply controlled breadth over structures and mechanisms; semantic operational SCMs provide human-auditable, semantically grounded environments beyond standard synthetic generators; and formula-grounded SCMs test discovery under explicit scientific mechanisms. Public real-world datasets provide an additional external-validity check. Experiments across classical, neural, and pretrained methods reveal substantial ranking shifts across SCM families and protocols, showing that strong performance in one benchmark regime does not reliably transfer to others. These results highlight benchmark diversity and pretraining--evaluation overlap as central challenges for evaluating causal discovery in the foundation model era.