混合图模型的贝叶斯后验学习
Bayesian Posterior Learning of Mixed Graphical Models
- University College London(伦敦大学学院)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出Mixed Graph WWA方法,通过潜在高斯变量和两种似然设定实现混合图模型的贝叶斯结构学习,在模拟和乳腺癌数据上验证了其准确性和高效性。
中文摘要 AI 辅助
混合图模型(MGMs)通过同时处理连续变量和离散变量集合,为从异构数据中进行结构学习提供了一个灵活的框架。由于图空间探索和后验计算的组合复杂性,MGMs的贝叶斯推断仍然具有挑战性。在本文中,我们通过潜在高斯变量对离散分量进行建模,并考虑两种似然设定:基于copula的秩似然(产生copula-MGM)和基于截止点的probit公式(产生probit-MGM)。然后,我们提出了Mixed Graph WWA,一类用于贝叶斯MGM后验模拟的MCMC方法。基于WWA算法,我们开发了两种专门算法:用于copula-MGM的copula-WWA和用于probit-MGM的probit-WWA。这两种方法都利用潜在高斯表示,通过吉布斯采样方案进行后验推断,该方案在潜在变量增广和图结构更新之间交替进行。通过广泛的模拟研究,我们证明了所提出的方法在图形恢复准确性方面达到或优于现有方法(包括基于Birth-Death MCMC方法的copula-BD MCMC和probit-BD MCMC),同时保持高效的后验探索和每单位计算时间的有效样本量。我们进一步通过将其应用于PAM$50$乳腺癌基因表达数据集来说明我们方法的实用性,其中推断出的MGM揭示了基因表达谱与癌症亚型之间的有意义依赖关系。这些结果突出了Mixed Graph WWA方法作为MGM中贝叶斯结构学习的可扩展且原理性工具的有效性。
英文摘要
Mixed Graphical Models (MGMs) provide a flexible framework for structure learning from heterogeneous data by treating sets of both continuous and discrete. Bayesian inference for MGMs remains challenging due to the combinatorial complexity of graph space exploration and posterior computation. In this paper, we model discrete components through latent Gaussian variables and consider two likelihood specifications: a copula-based ranked likelihood, yielding the copula-MGM, and a probit formulation based on cut-off points, yielding the probit-MGM. We then propose the Mixed Graph WWA, a class of MCMC methods for posterior simulation in Bayesian MGMs. Building upon the WWA algorithm, we develop two specialized algorithms: copula-WWA for copula-MGMs and probit-WWA for probit-MGMs. Both methods exploit the latent Gaussian representations to perform posterior inference through a Gibbs sampling scheme that alternates between latent-variable augmentation and graph-structure updates. Through extensive simulation studies we demonstrate that the proposed methods achieve graph recovery accuracy comparable to or better than existing approaches, including copula-BD MCMC and probit-BD MCMC based on the Birth-Death MCMC methodology, while maintaining efficient posterior exploration and favorable effective sample size per unit computational time. We further illustrate the practical utility of our approach through an application to the PAM$50$ breast cancer gene expression dataset, where the inferred MGMs reveal meaningful dependencies between gene expression profiles and cancer subtypes. These results highlight the effectiveness of the Mixed Graph WWA method as a scalable and principled tool for Bayesian structure learning in MGMs.