arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

映射并推进非线性因果发现的可扩展性-准确性前沿

Mapping and Advancing the Scalability-Accuracy Frontier of Nonlinear Causal Discovery

Hendrik Suhr, Sascha Xu, Jilles Vreeken

arXiv 2610.03258首次发表:更新:

发表机构

CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究比较四种非线性因果发现方法,揭示各自瓶颈,并提出SPADE评分方案,通过重用充分统计量将组合搜索复杂度从O(nd^3)降至O(nd^2+d^3),在保持高准确性的同时大幅提升可扩展性。

AI 中文摘要

可扩展的非线性因果发现需要将灵活的机制估计器与在大图空间上的高效搜索相结合的方法。为应对这一挑战,已有多个算法族被提出,然而它们的准确性与运行时间之间的权衡仍未被充分理解。我们实证比较了四种主要方法:可微结构学习、摊销结构学习、分数匹配和组合搜索。我们的结果揭示了互补的瓶颈:可微和摊销方法扩展性好但存在准确性差距,分数匹配方法在低维度下可能准确但随特征规模增加而迅速退化,组合方法保持准确但因重复和冗余的局部评分而变慢。受此瓶颈启发,我们开发了SPADE,一种基于样条的评分评估方案,该方案编译一次充分统计量并在整个组合搜索中重用它们。在有界入度下,其高斯变体将算法复杂度从O(nd^3)降低到O(nd^2+d^3)。实证上,SPADE将观察到的可扩展性-准确性前沿移动了数量级:它在几秒内解决具有160K样本的100变量问题,在几分钟内解决具有2.5K样本的1600变量问题,同时在合成和真实世界基准上保持高结构准确性。这些结果揭示了组合搜索实际规模的显著转变,并强调了沿完整准确性-运行时间前沿评估可扩展因果发现方法的重要性。

英文摘要

Scalable nonlinear causal discovery requires methods that combine flexible mechanism estimators with efficient search over large graph spaces. Several algorithmic families have been proposed to address this challenge, yet their accuracy-runtime trade-offs remain poorly understood. We empirically compare the four major approaches: differentiable structure learning, amortized structure learning, score-matching, and combinatorial search. Our results reveal complementary bottlenecks: differentiable and amortized methods scale well but exhibit an accuracy gap, score-matching methods can be accurate in low dimensions but degrade quickly for increasing feature sizes, and combinatorial methods remain accurate but are slowed by repeated and redundant local scoring. Motivated by this bottleneck, we develop SPADE, a spline-based score-evaluation scheme that compiles sufficient statistics once and reuses them throughout combinatorial search. Under bounded indegree, its Gaussian variant reduces algorithmic complexity from O(nd^3) to O(nd^2+d^3). Empirically, SPADE shifts the observed scalability-accuracy frontier by orders of magnitude: it solves 100-variable problems with 160K samples in seconds and 1600-variable problems with 2.5K samples in minutes, while retaining high structural accuracy across synthetic and real-world benchmarks. These results reveal a substantial shift in the practical scale of combinatorial search and highlight the importance of evaluating scalable causal-discovery methods along the full accuracy-runtime frontier.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑