arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不精确匹配何时能在无需调整的情况下确保平衡性与推断有效性?

When Does Inexact Matching Ensure Balance and Inference without Adjustment?

Ying Jin

arXiv 2610.10873首次发表:更新:

发表机构

University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对d维连续协变量的一对一匹配,证明d≤3时匹配设计可保证平衡性并支持多种有效推断,d≥4时平衡性失效,且通过数值实验验证了理论结果。

AI 中文摘要

无放回的一对一匹配是观察性研究设计中构建可比处理组与对照组样本的经典方法,它将每个处理单元与一个不同的控制单元配对,同时最小化协变量距离目标。对于连续协变量,匹配后的配对通常仍不精确,这会导致后续分析出现偏差。其关键理论性质,如匹配配对间的不平衡程度,以及何时该不平衡可忽略以支持有效推断,目前仍不明确。本文分析基于d维连续协变量且采用二次协变量距离目标的一对一匹配。首先,研究发现当d≤3时,在倾向得分满足每个处理样本附近有充足控制样本的标准条件下,对于具有共同一阶和二阶导数界的光滑函数族,不平衡(组内均值之差)在根n尺度上是可忽略的;但这种平衡性受维度限制,当d=4时不平衡在根n尺度上不可忽略,当d>4时不平衡会主导根n速率,我们构造了此类示例。其次,证明当d≤3时,该匹配设计可对处理组平均处理效应、分布处理效应及分位数处理效应进行有效的Wald型和自助法推断,因此这种不依赖结果的匹配设计可支持多种后续推断,无需针对目标调整设计。最后,基于该匹配设计的配对随机化推断在超总体意义下对d≤3渐近有效,但当d=4时可能失效。我们通过数值实验验证了这些理论结果。

英文摘要

One-to-one matching without replacement is a classical approach to constructing comparable treated and control samples in the design of observational studies. It pairs each treated unit with a distinct control while minimizing a covariate distance objective. With continuous covariates, the matched pairs generally remain inexact, which contributes to bias in downstream analysis. Its key theoretical properties, such as the resulting imbalance between matched pairs and when it is negligible to support valid inference, remain unclear. In this paper, we analyze one-to-one matching based on $d$-dimensional, continuous covariates with a quadratic covariate-distance objective. First, we find that when $d\leq 3$, under standard conditions on the propensity score ensuring abundant control samples near each treated sample, the imbalance (difference between within-group averages) is root-$n$ negligible uniformly over the family of smooth functions with a common first- and second-order derivative bound. However, such balance is subject to a dimension restriction, as we construct examples in which the imbalance is root-$n$ non-negligible when $d=4$ and dominates root-$n$ rate when $d>4$. Second, we show that when $d\leq 3$, the matched design allows valid Wald-type and bootstrap inference for the average treatment effect on the treated, distributional treatment effects, and quantile treatment effects. Thus, the same outcome-blind matched design supports various downstream inferences without having to tailor the design to the targets. Finally, paired randomization inference based on the matched design is asymptotically valid in the super-population sense for $d\leq 3$ but can fail when $d=4$. We corroborate the theoretical results with numerical experiments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑