arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18417stat.ME

质心参考马氏匹配(CRM):面向大规模观察性研究因果推断的可扩展、基于表示的框架

Centroid-Referenced Mahalanobis Matching (CRM): A Scalable, Representation-Based Framework for Causal Inference in Large Observational Studies

Keming Hu, Yingpei He

AI总结:

针对大规模观察性研究因果推断的匹配方法存在成本高、重叠有限时改变目标总体的问题,本文提出CRM框架,其计算成本低、保留处理单元比例高、平衡表现优且速度快,贡献在于可扩展性与显式支撑诊断。

AI中文摘要:

因果推断中的匹配方法在大规模场景下计算成本高昂,且在重叠性有限时会悄悄改变目标总体。我们提出质心参考马氏匹配(CRM),该方法将全局成对搜索替换为在两个参考坐标中的分层抽样:每个单元到处理组质心的马氏距离,以及其沿处理-对照均值偏移的费希尔坐标。所有协变量均通过处理组协方差几何输入;因此CRM并非主成分预处理后再进行近邻匹配。对于$n$个单元和$p$个预处理协变量,其实现的成本为$O(np^2+p^3+n\log n)$,当$n \ge p$时简化为$O(np^2+n\n\log n)$。我们推导了误差分解,将其分为表示、支撑、离散化和随机成分。匹配前短缺分数$\hat{\pi}$估计总体支撑限制$\pi$,该值在处理效应异质性有界时纳入间隙界。最终保留率按容量驱动的排除项单独报告。在表示充分性、平滑性和单元容量充足的条件下,CRM具有保守的二维直方图均方误差界$O(n_T^{-1/2})$;表示充分性是额外假设,而非给定原始协变量下可忽略性的结果。在Criteo数据集上,CRM保留至少99.4%的处理单元,在36种大规模配置中,其最大标准化均值差(MaxSMD)低于校正倾向得分匹配的情况有31种,且速度快约一个数量级。中等规模模拟显示,部分成对和加权基线在平衡性上更优,这明确了CRM的贡献在于可扩展性和显式支撑诊断,而非通用有限样本优势。

英文摘要:

Matching for causal inference can be computationally expensive at scale and can silently change the target population when overlap is limited. We propose Centroid-Referenced Mahalanobis Matching (CRM), which replaces global pairwise search with stratified sampling in two reference coordinates: each unit's Mahalanobis distance from the treated centroid and its Fisher coordinate along the treated-control mean shift. All covariates enter through the treated covariance geometry; CRM is therefore not principal-component preprocessing followed by nearest-neighbor matching. For $n$ units and $p$ pretreatment covariates, its implemented cost is $O(np^2+p^3+n\log n)$, simplifying to $O(np^2+n\log n)$ when $n \ge p$. We derive an error decomposition separating representation, support, discretization, and stochastic components. A pre-matching shortage fraction $\hatπ$ estimates the population support restriction $π$, which enters a gap bound under bounded treatment-effect heterogeneity. Final retention is reported separately for capacity-driven exclusions. Under representation sufficiency, smoothness, and adequate cell capacity, CRM has a conservative two-dimensional histogram mean-squared-error bound $O(n_T^{-1/2})$; representation sufficiency is an additional assumption, not a consequence of ignorability given the original covariates. On Criteo, CRM retains at least 99.4% of treated units, has lower MaxSMD than corrected propensity-score matching in 31 of 36 large-scale configurations, and is roughly an order of magnitude faster. Moderate-size simulations favor some pairwise and weighting baselines on balance, locating CRM's contribution in scalability and explicit support diagnostics rather than universal finite-sample dominance.

补充信息

↑