发表机构
University of California, Los Angeles; University of California, San Diego; Xiamen University(加州大学洛杉矶分校; 加州大学圣地亚哥分校; 厦门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对合成表格数据生成器仅优化分布保真度无法保留因果效应的问题,提出因果惩罚的TabDDPM框架,经实验验证可在提升因果保真度的同时保持竞争力统计保真度。
AI 中文摘要
合成表格数据生成器通常以分布保真度为优化目标,但仅统计相似性无法保证因果效应的保留。本文研究能否在全生成式表格模型中直接提升因果保真度,将因果保真度定义为真实数据与合成数据推断分布的差异,理论结果表明高统计保真度通常不意味着高因果保真度。我们提出因果感知训练框架,在生成目标中加入因果差异惩罚,实例化为因果惩罚的TabDDPM,采用on-policy得分函数估计器优化,还建立了因果正则化提升期望因果保真度的条件。在多种处理效应模拟及两个基准数据集上的实验,评估了该方法在提升因果保真度的同时保持竞争力统计保真度的能力。
英文摘要
Synthetic tabular generators are commonly optimized for distributional fidelity, but statistical similarity alone does not guarantee preservation of causal effects. In this paper, we study whether causal fidelity can be improved directly within a fully generative tabular model. Causal Fidelity is defined with respect to a target estimand as the discrepancy between inferential distributions obtained from real and synthetic data, and theoretical results show that high statistical fidelity does not generally imply high causal fidelity. We then propose a causal-fidelity-aware training framework which adds a causal discrepancy penalty to the generative objective. The framework is instantiated with a causal-penalized TabDDPM and optimized using an on-policy score-function estimator. We further establish conditions under which causal regularization improves expected causal fidelity. Experiments across diverse treatment-effect simulations and two benchmark datasets evaluate the ability of our method to improve causal fidelity while preserving competitive statistical fidelity.