arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03868cs.LGcs.AI

GENESIS:迈向可解释的因果发现

GENESIS: Towards Explainable Causal Discovery

Abhinav Thorat, Ravi Kumar Kolla, Vishak K Bhat, Harsh Vardhan Singh Chauhan, Niranjan Pedanekar

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出可解释混合因果发现框架GENESIS,实现100%决策可追溯性,在多数基准数据集上性能优于纯统计方法,可与先进LLM辅助方法媲美。

中文摘要 AI 辅助

从观测数据进行因果发现(CD)面临两个基本挑战:其一,纯统计方法在小样本场景下往往缺乏解决结构歧义的能力;其二,尽管大语言模型(LLM)辅助的混合方法可通过语义推理提升结构恢复效果,但该推理对单个边决策的影响仍极不透明。因此,现有混合方法无法满足一项基本要求:解释为何学习到的有向无环图(DAG)中包含或排除某条特定边,这在无真实DAG的实际应用中至关重要,因为每项结构决策都必须得到独立验证。我们将该要求形式化为决策可追溯性,要求每条推断边都有可审计的统计证据、马尔可夫毯一致性或明确的领域推理作为支撑。我们提出GENESIS,这是一个可解释的混合CD框架,它将图构建分解为可解释的决策点:GENESIS首先识别并对三节点结构基序(包括链、叉和对撞机)进行评分,以建立透明的结构先验,随后通过将这些先验与观测证据相整合来逐步优化图,仅在统计证据不足时调用领域知识。根据设计,每条边决策都通过可审计的证据来源得以解决。实验表明,GENESIS在所有场景下均实现了100%的决策可追溯性,将可解释性确立为因果发现的核心目标;尽管增加了该要求,GENESIS在所有样本场景下的多数基准数据集上,结构汉明距离(SHD)指标始终优于纯统计CD方法,且性能可与最先进的LLM辅助方法相媲美。

英文摘要

Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve structural ambiguities in low-sample regimes. Second, although LLM-assisted hybrid approaches improve structure recovery through semantic reasoning, the influence of that reasoning on individual edge decisions remains largely opaque. Consequently, existing hybrid methods fail to satisfy a fundamental requirement: explaining why a particular edge is included or excluded in the learned directed acyclic graph (DAG). This is critical in real-world applications, where no ground-truth DAG exists and every structural decision must be independently justified. We formalize this requirement as decision traceability, requiring every inferred edge to be supported by auditable statistical evidence, Markov Blanket consistency, or explicit domain reasoning. We propose GENESIS, an explainable hybrid CD framework that decomposes graph construction into interpretable decision points. GENESIS first identifies and scores three-node structural motifs, including chains, forks, and colliders, to establish transparent structural priors, then progressively refines the graph by integrating these priors with observational evidence, invoking domain knowledge only when statistical evidence is insufficient. By design, every edge decision is resolved through an auditable source of evidence. Experiments show that GENESIS achieves 100% decision traceability across all settings, establishing explainability as a first-class objective in causal discovery. Despite this additional requirement, GENESIS consistently outperforms purely statistical CD methods on the majority of benchmark datasets across all sample regimes in terms of Structural Hamming Distance (SHD), while achieving performance comparable to state-of-the-art LLM-assisted approaches.

补充信息

↑