发表机构
School of Statistics, East China Normal University; Department of Statistics, London School of Economics and Political Science(华东师范大学统计学院; 伦敦政治经济学院统计学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出DAG-CLIP框架,用于存在潜变量时学习有向无环图,通过两步复合似然算法筛选并剪枝结构,恢复真实DAG的马尔可夫等价类,并在模拟及教育调查数据上验证有效性。
AI 中文摘要
学习有向无环图(DAG)以揭示因果机制已引起机器学习领域的广泛关注。虽然大多数现有方法仅关注观测变量,但许多具有实质意义的变量是由统计测量模型定义的潜在构念。这类构念在社会科学和行为科学中尤为常见。在本文中,我们提出了一个通用的统计框架,用于在图的部分或全部节点为潜变量时进行DAG学习,该框架在测量模型中容纳非高斯显变量,如二分类和分类数据。为了克服标准边际似然中高维积分的计算负担,我们开发了基于复合似然筛选和迭代剪枝的DAG学习算法(DAG-CLIP),这是一种两步的基于复合似然的学习算法。该算法首先求解一个平滑的无环约束优化问题,以筛选出错误设定的DAG结构,随后在马尔可夫等价类(MEC)空间内进行基于BIC准则的复合似然向后删除,以识别稀疏的DAG。我们证明了该算法在恢复真实DAG的MEC方面具有统计一致性,并通过广泛的模拟研究和一项针对大规模教育调查数据的实际应用展示了其有效性。
英文摘要
Learning directed acyclic graphs (DAGs) to uncover causal mechanisms has attracted substantial attention in machine learning. While most existing methods focus exclusively on observed variables, many variables of substantive interest are latent constructs defined by statistical measurement models. Such constructs are particularly common in the social and behavioral sciences. In this paper, we propose a general statistical framework for DAG learning when some or all nodes of the graph are latent variables, accommodating non-Gaussian manifest variables such as binary and categorical data in the measurement model. To overcome the computational burden of high-dimensional integration in standard marginal likelihoods, we develop DAG learning via Composite-Likelihood-based screening and Iterative Pruning (DAG-CLIP), a two-step composite-likelihood-based learning algorithm. This algorithm first solves a smooth acyclicity-constrained optimization problem to screen out misspecified DAG structures, and subsequently performs BIC-guided composite-likelihood backward deletion within the Markov equivalence class (MEC) space to identify a sparse DAG. We establish the statistical consistency of the algorithm in recovering the MEC of the true DAG, and demonstrate its effectiveness through extensive simulation studies and a real-world application to large-scale educational survey data.