发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究从干预数据学习因果图的难题,提出扩展因子图的摊销贝叶斯因果发现方法(ABCDEFG),可处理大规模图及未知干预目标,保证无环性,估计后验分布,在模拟和实际数据上表现优异,能识别基因靶点。
AI 中文摘要
从干预数据中学习因果图是一个具有广泛应用的挑战性问题。例如,在分子生物学中,核心目标是从大规模扰动数据中揭示基因调控网络。理想算法应能扩展到数千个节点,纳入目标未知的干预,量化不确定性并提供可识别性保证。现有方法常无法满足这些标准。为此我们开发了扩展因子图的摊销贝叶斯因果发现(ABCDEFG)方法。该方法保证精确无环性,可扩展到数千节点的图,能自然处理目标未知的干预。此外,ABCDEFG估计后验分布,其最大后验估计可证明识别出直至等价类的真实因果图。在模拟数据集上,ABCDEFG达到了当前最优精度,应用于大规模单细胞扰动数据时,能识别生长因子已确立和新的基因靶点。
英文摘要
Learning causal graphs from interventional data is a challenging problem with broad applications. In molecular biology, for example, a central goal is to uncover gene regulatory networks from large-scale perturbation data. An ideal algorithm for this task should scale to thousands of nodes, incorporate interventions even when their targets are unknown, quantify uncertainty, and provide identifiability guarantees. However, existing approaches---e.g. approaches using score-based optimization or approximate Bayesian inference---often fail to meet all of these criteria. To address these limitations, we develop Amortized Bayesian Causal Discovery of Extended Factor Graphs (ABCDEFG). Our method guarantees exact acyclicity, scales to graphs with thousands of nodes, and naturally handles interventions even when their targets are unknown. Additionally, ABCDEFG estimates a posterior distribution whose maximum a posteriori estimate provably identifies the true causal graph up to an equivalence class. On simulated datasets, ABCDEFG achieves state-of-the-art accuracy, producing a well-calibrated posterior distribution while outperforming previous score-based and approximate Bayesian methods. Applied to large-scale single-cell perturbation data, ABCDEFG identifies both established and novel gene targets of growth factors.