arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从不完整数据中通过观测数据似然学习因果归一化流

Learning Causal Normalizing Flows from Incomplete Data via Observed-Data Likelihood

Trung-Dung Hoang, Alceu Bissoto, Tim Flühmann, David Herzig, Christos Nakas, Lia Bally, Lisa M. Koch

arXiv 2609.37664首次发表:更新:

发表机构

University of Bern; Inselspital, Bern University Hospital; Diabetes Center Berne; University of Thessaly(伯尔尼大学; 伯尔尼大学医院; 伯尔尼糖尿病中心; 色萨利大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MissCNF,通过最大化部分观测样本的边际似然直接在不完整数据上训练因果归一化流,无需丢弃数据或插补,在多种缺失机制下优于现有方法。

AI 中文摘要

因果归一化流(CNFs)能够在给定因果结构的情况下,从观测数据中进行因果推断,但它们假设训练数据是完全观测的。我们提出了MissCNF,它通过最大化每个部分观测样本的边际似然,直接在不完整数据上训练CNFs,而无需丢弃行或构建完整数据集。得益于CNFs自回归分解中编码的因果结构,只有观测集合的祖先闭包中的缺失变量被积分掉,而其他变量则无需计算即可丢弃。我们进一步建立了MissCNF恢复真实联合分布的条件,并引入了“因果族正性”(causal-family positivity),即使在数据集中没有任何记录是完整的情况下,识别也是可能的。我们将MissCNF与两种常见的缺失数据处理策略进行比较:列表删除和插补后拟合流程。在八个合成因果基准、三种缺失机制以及缺失率高达90%的情况下,MissCNF在24个非线性MCAR和MAR设置中的23个以及所有非线性MNAR设置中取得了最低的KL散度,并在24个设置中的20个中取得了最低的反事实误差。在线性SCM上,线性插补表现最佳,MissCNF在24个设置中的22个中排名前两位。

英文摘要

Causal Normalizing Flows (CNFs) enable causal inference from observational data given the causal structure, but they assume fully observed training data. We introduce MissCNF, which trains CNFs directly on incomplete data by maximizing the marginal likelihood of each partially observed sample, without discarding rows or constructing a completed dataset. Thanks to the causal structure encoded in the autoregressive factorization of CNFs, only missing variables in the ancestral closure of the observed set are integrated out, while the others are dropped without computation. We further establish the conditions under which MissCNF recovers the true joint distribution, and introduce \emph{causal-family positivity}, where identification is possible even when no record in the dataset is ever complete. We compare MissCNF with two common strategies for handling missing data: listwise deletion and impute-then-fit pipelines. Across eight synthetic causal benchmarks, three missingness mechanisms, and missing rates up to $90\%$, MissCNF achieves the lowest KL divergence in 23 of 24 nonlinear MCAR and MAR settings and in all nonlinear MNAR settings, as well as the lowest counterfactual error in 20 of 24 settings. On linear SCMs, where linear imputation performs best, MissCNF ranks in the top two in 22 of 24 settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑