arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38618cs.LGcs.AI

具有潜在混杂因子的循环线性高斯模型的可微结构学习

Differentiable Structure Learning for Cyclic Linear Gaussian Models with Latent Confounders

  • EPFL(洛桑联邦理工学院)
  • Leiden University(莱顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Sadegh Khorasani, Ali Najar, Saber Salehkaleybar, Negar Kiyavash

AI总结:

本文提出一种可微结构学习方法,用于含潜在混杂因子的循环线性高斯模型,通过伯努利门和闭式可微惩罚实现全局一致的结构学习,实验显示恢复误差更低。

AI中文摘要:

我们研究在存在有向循环和给定最大数量上界的未知外生潜在混杂因子的情况下,从观测数据中进行线性高斯结构因果模型中的因果结构学习。我们推导了观测变量的协方差,并引入了边际拟等价性,该性质刻画了不同因果模型何时共享它们所能生成的观测分布的一个全维子集。我们将结构学习表述为高斯负对数似然的最小化,并带有一个对数尺度的复杂度惩罚项,该惩罚项对直接边和潜在变量进行计数。对于固定数量的观测变量和固定的潜在变量上界,我们在代数忠实性、结构最小性和模型重叠假设下,建立了全局分数最小化器在边际拟等价性下的一致性。我们使用伯努利门对直接边和候选潜在变量的包含进行参数化,其连续概率与结构系数联合优化。对这些门上的惩罚负对数似然进行平均,得到一个具有闭式可微复杂度惩罚的目标函数。我们证明了该期望目标与相应的离散结构学习目标具有相同的全局下确界。实验结果表明,在几种实验设置中,我们的方法比先前方法实现了更低的恢复误差。

英文摘要:

We study causal structure learning from observational data in linear Gaussian structural causal models in the presence of directed cycles and an unknown number of exogenous latent confounders, bounded by a given maximum. We derive the covariance of the observed variables and introduce marginal quasi-equivalence, which characterizes when different causal models share a full-dimensional subset of the observational distributions they can generate. We formulate structure learning as minimization of the Gaussian negative log-likelihood with a logarithmically scaled complexity penalty that counts directed edges and latent variables. For a fixed number of observed variables and a fixed upper bound on latent variables, we establish consistency of global score minimizers up to marginal quasi-equivalence under algebraic faithfulness, structural minimality, and model-overlap assumptions. We parameterize the inclusion of directed edges and candidate latent variables using Bernoulli gates, whose continuous probabilities are optimized jointly with the structural coefficients. Averaging the penalized negative log-likelihood over these gates yields an objective with a closed-form differentiable complexity penalty. We prove that this expected objective has the same global infimum as the corresponding discrete structure-learning objective. Experimental results show that our approach achieves lower recovery error than previous methods in several experimental settings.

补充信息

↑