发表机构
Department of Statistics, Korea University, Seoul, Republic of Korea(统计系,韩国大学,首尔,大韩民国)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究观察性因果分析中数据中毒问题,提出针对增强逆概率加权估计的数据中毒审计方法,通过指定相关参数、进行贪婪扫描等操作,推导总影响分数和有限预算界限,经模拟验证可支持可靠因果报告及源级保障设计。
AI 中文摘要
观察性因果分析越来越多地跨站点、供应商和收集系统汇总记录,容易受到仅追加攻击,即通过策略性选择似是而非的记录来改变报告的处理效应。我们为增强逆概率加权估计开发了一种数据中毒审计。分析师指定可行记录的有限目录、追加预算和嵌套源容量,对手选择可行子集以最大化在指定方向上的变动。在预处理和干扰拟合固定的情况下,我们提出一种贪婪扫描,计算每个追加预算下的确切有限样本最坏情况变动。为考虑干扰重新拟合,我们推导了一个总影响分数,结合每条记录的直接贡献及其通过倾向和结果模型的效应。我们还为完全重新拟合的估计获得了一个保守的有限预算界限。大量模拟验证了确切结果,并表明总影响改善了局部重新拟合预测,而多站点和公共数据分析表明在小追加预算下存在实质性敏感性。通过将对抗性数据构成风险转化为变动曲线和关键预算,该框架支持更可靠的因果报告和源级保障设计。
英文摘要
Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect. We develop a data-poisoning audit for augmented inverse-probability-weighted estimation. The analyst specifies a finite catalog of feasible records, an append budget, and nested source capacities, and the adversary selects a feasible subset to maximize movement in a prespecified direction. With preprocessing and nuisance fits held fixed, we propose a greedy scan that computes the exact finite-sample worst-case movement at every append budget. To account for nuisance refitting, we go on to derive a total-influence score combining each record's direct contribution with its effect through the propensity and outcome models. We further obtain a conservative finite-budget bound for the fully refitted estimate. Extensive simulations validate the exact result and show that total influence improves local refit prediction, while multisite and public-data analyses demonstrate material sensitivity at small append budgets. By translating adversarial data-composition risk into movement curves and critical budgets, the framework supports more reliable causal reporting and the design of source-level safeguards.