arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非参数正态贝叶斯学习有向无环图:基于Gamma和逆Gamma创新先验的闭式得分与知情采样

Nonparanormal Bayesian Learning of Directed Acyclic Graphs under Gamma and Inverse-Gamma Innovation Priors: Closed-Form Scores and Informed Sampling

Samaneh Nazari, Mohammad Arashi

arXiv 2609.13008首次发表:更新:

发表机构

Department of Statistics, Faculty of Mathematical Sciences, Ferdowsi University of Mashhad(马什哈德费尔多西大学数学科学学院统计学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对非高斯数据,提出非参数正态贝叶斯DAG学习方法,利用Gamma和逆Gamma先验获得闭式得分,构建知情采样器,在模拟和真实数据上优于高斯方法。

AI 中文摘要

有向无环图(DAG)的贝叶斯结构学习是重建生物网络的核心工具,但该方法通常假设数据服从联合高斯分布。在激励性的蛋白质组学应用中,这一假设常被偏态和重尾数据所违背,导致高斯DAG恢复出虚假或错误方向的边。我们开发了一个完全贝叶斯框架,用于非参数正态族中的DAG学习,用更弱的要求替代高斯性,即未知的严格递增边际变换是联合高斯的。基于潜在精度矩阵的修正Cholesky参数化,我们引入了两种创新方差先验:一种是非共轭的Normal-Gamma先验,它将系数收缩与方差正则化解耦;另一种是共轭的Normal-Inverse-Gamma先验。对于这两种先验,我们获得了节点边际似然的闭式表达式——对于Gamma先验通过第三类修正贝塞尔函数,对于逆Gamma先验通过Student-t形式——从而使得MCMC采样器的移动无需数值积分即可评分。利用这些,我们构建了一个局部平衡的知情采样器和一个无贝塞尔函数的得分,使采样器可扩展到数百个节点。在模拟数据上,当数据为高斯时,我们的非参数正态采样器与高斯方法相当,而当边际偏斜时则显著优于它们。在人类T细胞蛋白信号数据上,它们以高后验概率恢复了公认的相互作用,明显优于基于约束的竞争对手,并与高斯贝叶斯模型表现相当。

英文摘要

Bayesian structure learning for directed acyclic graphs (DAGs) is a central tool for reconstructing biological networks, yet it often assumes the data are jointly Gaussian. In motivating proteomic applications, this assumption is routinely violated by skewed and heavy-tailed data, causing Gaussian DAGs to recover spurious or misdirected edges. We develop a fully Bayesian framework for DAG learning in the nonparanormal family, replacing Gaussianity with the weaker requirement that unknown strictly increasing marginal transformations are jointly Gaussian. Working on the modified Cholesky parameterization of the latent precision matrix, we introduce two innovation-variance priors: a non-conjugate Normal-Gamma prior, which decouples coefficient shrinkage from variance regularization, and a conjugate Normal-Inverse-Gamma prior. For both, we obtain the node-wise marginal likelihood in closed form -- through a modified Bessel function of the third kind for the Gamma prior and a Student-$t$ form for the Inverse-Gamma prior -- allowing MCMC sampler moves to be scored without numerical integration. Exploiting these, we build a locally-balanced informed sampler and a Bessel-free score that scales the sampler to hundreds of nodes. On simulated data, our nonparanormal samplers match Gaussian methods when data are Gaussian and dominate them sharply when margins are skewed. On human T-cell protein-signalling data, they recover well-established interactions at high posterior probability, clearly outperforming constraint-based competitors and performing comparably to a Gaussian Bayesian model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑