发表机构
Vanderbilt University Medical Center; King’s College London(范德比尔特大学医学中心; 伦敦国王学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对因果推断的差分隐私合成数据,提出因果工作负载,围绕双稳健因果估计器的正交矩设计DP查询集,经最大熵校准重建合成数据,还引入相关工具,能在不额外隐私支出下支持多种分析,且揭示了分布保真度与因果推断的权衡。
AI 中文摘要
基于工作负载的差分隐私(DP)合成数据方法私下测量聚合查询,并将有噪声的答案后处理为合成记录。通用工作负载可实现强大的分布保真度,但诸如平均治疗效果(ATE)等因果估计量取决于治疗组平衡和结果矩,而通用边缘分布无需保留这些。我们提出了因果工作负载:围绕双稳健因果估计器使用的正交矩设计的DP查询集。发布的工作负载可由稳定矩图估计器直接使用,或通过最大熵校准重建为可重复使用的合成数据;我们的理论将ATE误差分解为采样、隐私、工作负载近似、蒙特卡罗和校准项。我们还引入了Causal - AIM,一种自适应工作负载选择器,以及用于从DP合成数据获取置信区间的噪声感知多重插补(NA + MI)过程。由于工作负载只发布一次,相同的DP合成表可支持ATE、ATT和子组分析,而无需额外的隐私支出。从经验上看,因果工作负载在严格的隐私预算和校准不确定性方面最有用,而随着隐私放宽,通用工作负载在点RMSE方面通常具有优势。更广泛的教训是一种权衡:分布保真度有助于点精度,但有效的因果推断需要保留因果矩并传播DP噪声,而不是将合成行视为真实的。
英文摘要
Workload-based differentially private (DP) synthetic data methods privately measure aggregate queries and post-process the noisy answers into synthetic records. Generic workloads can achieve strong distributional fidelity, but causal estimands such as the average treatment effect (ATE) depend on treatment-arm balance and outcome moments that generic marginals need not preserve. We propose causal workloads: DP query sets designed around the orthogonal moments used by doubly robust causal estimators. The released workload can be used directly by stable moment-map estimators or reconstructed by maximum-entropy calibration into reusable synthetic data; our theory decomposes ATE error into sampling, privacy, workload-approximation, Monte Carlo, and calibration terms. We also introduce Causal-AIM, an adaptive workload selector, and a noise-aware multiple-imputation (NA+MI) procedure for confidence intervals from DP synthetic data. Because the workload is released once, the same DP synthetic table can support ATE, ATT, and subgroup analyses without additional privacy spending. Empirically, causal workloads are most useful at strict privacy budgets and for calibrated uncertainty, while generic workloads often retain an advantage for point RMSE as privacy relaxes. The broader lesson is a tradeoff: distributional fidelity can help point accuracy, but valid causal inference requires preserving causal moments and propagating DP noise rather than treating synthetic rows as real.
CommentsAccepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026). Includes appendices. Code: https://github.com/AsiaeeLab/causal-aim References now show missing arxiv ids