实验设计的局限性:超越低维的协变量平衡
The Limits of Experimental Design: Covariate Balance Beyond Low Dimension
浏览论文内容
中文总结 AI 辅助
该研究证明低维协变量平衡的实验设计存在局限性,提出基于偏差最小化的新设计可处理高维协变量,结合匹配后在模拟中能降低处理效应估计的方差。
中文摘要 AI 辅助
我们研究在有限样本下,实验设计以未调整处理效应估计的额外方差衡量,趋近半参数效率界的速度。我们证明了一个不可能性定理:在弱条件下,除非协变量维度$d \ll \log n$,否则没有设计能在平滑结果模型上一致地趋近方差界。即使在具有数千个单元的实验中,这也仅允许少量协变量。受此启发,我们提出了基于偏差最小化的新设计,这些设计试图控制受限复杂度非参数函数类上的不平衡。此类设计能以快速速率达到其对应的受限效率目标,在加性非参数设定中允许$d \ll n$个协变量。它们还可与匹配结合,以防范未建模的结果变异。在针对12个已发表实验校准的模拟中,我们的设计在所有实证场景下都比配对随机化匹配减少了方差。
英文摘要
We study how fast experimental designs can approach the semiparametric efficiency bound in finite samples, as measured by the excess variance of unadjusted treatment effect estimation. We prove an impossibility theorem: under weak conditions, no design can approach the variance bound uniformly over smooth outcome models unless covariate dimension $d \ll \log n$. Even in experiments with thousands of units, this permits only a handful of covariates. Motivated by this, we propose new designs based on discrepancy minimization that instead attempt to control imbalances over restricted-complexity nonparametric function classes. Such designs achieve fast rates to their corresponding restricted efficiency targets, permitting $d \ll n$ covariates in an additive nonparametric specification. They can also be combined with matching to protect against unmodeled outcome variation. In simulations calibrated to 12 published experiments, our designs reduce variance relative to matched pairs randomization in every empirical setting.