arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAFE:通过快速后验估计进行反事实预测

CAFE: Counterfactual Prediction via Fast Posterior Estimation

Xinyan Han, Xiaoyu Lin, Hao Zou, Xingxuan Zhang, Bo Li, Peng Cui

arXiv 2609.32167首次发表:更新:

AI 中文总结

CAFE提出一种基于Transformer的摊销推断框架,通过快速后验估计直接近似贝叶斯反事实后验预测分布,在可识别及结构不确定场景下均表现准确,并验证了现实应用中的鲁棒性。

AI 中文摘要

反事实预测根据个体的事实观测,估计其在替代干预下的结果。在没有额外假设的情况下,此类结果通常无法仅从观测数据中识别。即使在完全观测的加性噪声模型(ANMs)类别中,不同的因果图也能生成相同的观测分布,却可能意味着不同的个体反事实结果。基于单一估计图的预测忽略了这种结构不确定性。因此,我们针对贝叶斯反事实后验预测分布进行建模,该分布结合了来自合理SCM的预测。我们提出了CAFE(通过快速后验估计进行反事实预测),这是一种摊销推断框架,直接近似贝叶斯反事实后验预测分布。我们在由ANMs上的多样化先验生成的合成反事实任务上预训练了一个基于Transformer的模型。给定一个观测数据集、一个个体的事实观测和一个干预,CAFE在单次前向传播中近似相应的后验预测分布。实验表明,CAFE在可识别设置中准确预测个体反事实结果,并在存在由观测上不可区分的因果图引起的结构不确定性时近似后验预测分布。在现实制造和葡萄栽培环境中的强劲表现进一步证明了其超越训练先验假设的经验鲁棒性。

英文摘要

Counterfactual prediction estimates an individual's outcome under an alternative intervention given their factual observations. Such outcomes are generally not identifiable from observational data without additional assumptions. Even within the class of fully observed additive noise models (ANMs), different causal graphs can generate the same observational distribution yet imply different individual counterfactual outcomes. Predictions based on a single estimated graph ignore this structural uncertainty. We therefore target a Bayesian counterfactual posterior predictive distribution that combines predictions from plausible SCMs. We introduce CAFE (\textbf{C}ounterf\textbf{A}ctual Prediction via \textbf{F}ast Posterior \textbf{E}stimation), an amortized inference framework that directly approximates the Bayesian counterfactual posterior predictive distribution. We pretrain a transformer-based model on synthetic counterfactual tasks generated from a diverse prior over ANMs. Given an observational dataset, an individual's factual observations, and an intervention, CAFE approximates the corresponding posterior predictive distribution in a single forward pass. Experiments show that CAFE accurately predicts individual counterfactual outcomes in identifiable settings and approximates the posterior predictive distribution when structural uncertainty induced by observationally indistinguishable causal graphs exists. Strong performance in realistic manufacturing and viticulture settings further demonstrates its empirical robustness beyond the assumptions of the training prior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑