发表机构
University of Waterloo; Vector Institute for Artificial Intelligence(滑铁卢大学; 向量人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对GP-SCM中离散结果的反事实推断,提出统一概率框架,推导精确噪声推断程序,并证明耦合误设会显著增加反事实误差。
AI 中文摘要
高斯过程结构因果模型(GP-SCMs)中的反事实推断主要针对连续内生变量开发,这限制了其在包含具有连续父节点的离散子节点的因果图中的应用。我们通过将GP预测器与显式外生噪声机制配对,引入了一个统一概率框架,用于处理异构变量类型的反事实推断。对于离散结果,我们使用二元变量的均匀阈值、名义类别的Gumbel-max竞争和有序变量的潜在高斯切点模型,推导出精确的条件噪声推断(abduction)程序。在每种情况下,我们通过干预传播推断出的噪声,同时考虑GP潜在函数中的后验不确定性,并证明所得机制再现了拟合模型的观测分布和干预分布。在具有已知真实反事实的合成SCM上,我们评估了估计准确性、一致性和对耦合误设的鲁棒性。一个关键发现是,将类别耦合应用于有序数据会使反事实误差大约增加三倍,即使观测拟合仍然相当,并且这种误差不会随着更多数据而减少。随着训练集的增长,拟合的结构方程收敛到真实值,而反事实误差则趋于平稳。相反,将虚假顺序强加于名义数据则会降低拟合方程本身的质量。因此,耦合的选择必须基于结构理由,而不能仅从拟合结果中读取。
英文摘要
Counterfactual inference in Gaussian-process structural causal models (GP-SCMs) has been developed primarily for continuous endogenous variables, limiting applicability to causal graphs that contain discrete child nodes with continuous parents. We introduce a unified probabilistic framework for counterfactual inference with heterogeneous variable types by pairing GP predictors with explicit exogenous noise mechanisms. For discrete outcomes, we derive exact conditional noise-abduction procedures using a uniform threshold for binary variables, a Gumbel-max race for nominal categories, and a latent Gaussian cut-point model for ordinal ones. In each case, we propagate abducted noise through interventions while accounting for posterior uncertainty in the GP latent functions, and prove that the resulting mechanisms reproduce the fitted model's observational and interventional distributions. On synthetic SCMs with known ground-truth counterfactuals, we evaluate estimation accuracy, consistency, and robustness to coupling misspecification. A key finding is that applying a categorical coupling to ordinal data inflates counterfactual error roughly threefold even when observational fit remains comparable, and that this error does not diminish with more data. As the training set grows, the fitted structural equation converges to the truth while the counterfactual error flattens onto a floor. In the reverse direction, forcing a false order onto nominal data instead degrades the fitted equation itself. The choice of coupling must therefore be justified on structural grounds rather than read off the fit.