发表机构
Technion Israel Institute of Technology(以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出OSIP框架,将观察性匹配建模为联合优化问题,通过两阶段算法在最大层数和精细平衡约束下实现高效平衡,并在基准数据集上验证其优越性。
AI 中文摘要
治疗组与对照组之间的协变量匹配仍然是观察性因果推断中的基础方法,旨在复制随机对照试验的协变量平衡。尽管方法论上取得了重大进展,应用研究人员仍面临在实现严格平衡与保留匹配样本统计功效之间的固有权衡。我们提出了一种名为最优研究内部分区(OSIP)的匹配方法的新框架。OSIP将观察性匹配视为一个联合优化问题,同时受限于层数上的最大基数约束和每层内匹配的严格精细平衡阈值。OSIP基于两阶段优化算法。在两个阶段中,它通过评估每个单元的估计倾向得分,将(高维)数据投影到一维空间。该方法基于将输入强制划分为估计倾向得分连续区间的子集。新颖之处在于我们将此任务视为全局优化问题。在第一阶段,我们针对辅助(一维)目标函数采用最优分区算法。在第二阶段,我们回到原始距离函数,并应用启发式方法对第一阶段的解进行局部优化。我们将我们的框架与文献中成熟的方法在广泛研究的经验基准上进行比较。我们使用标准的Lindner心血管数据集基准来证明OSIP为相应的多准则目标提供了高效的甜点解。这一结论也通过考虑其他基准数据集得到验证。
英文摘要
Covariate matching between treatment and control groups remains a foundational methodology in observational causal inference, designed to replicate the covariate balance of randomized controlled trials. Despite substantial methodological advancements, applied researchers continue to face an inherent trade-off between achieving rigorous balance and preserving the statistical power of the matched sample. We introduce a novel framework for a matching method named Optimal Study Interior Partitioning (OSIP). OSIP frames observational matching as a joint optimization problem governed simultaneously by a maximum cardinality constraint on the number of strata, and a strict fine balance threshold on the matching carried out in each strata. OSIP is based on two-stage optimization algorithm. In both stages it considers a projection of the (high-dimensional) data into one-dimensional space by evaluating an estimated propensity score for each unit. The method is based on enforcing a partition of the input to subsets of consecutive intervals of estimated propensity scores. The novelty arises because we consider this task as a global optimization problem. In the first stage, we employ an optimal partitioning algorithm with respect to an auxiliary (one-dimensional) goal function. In the second phase, we go back to the original distance function and apply heuristics approaches locally optimizing the solution from the first stage. We compare our framework with established methodologies from the literature across well-studied empirical benchmarks. We use the standard benchmark of Lindner cardiovascular dataset to demonstrate that OSIP provides a highly efficient sweet spot solution for the corresponding multi-criteria goal. This conclusion is verified by considering also other benchmark datasets.