arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2304.14545stat.MEcs.LGecon.EMstat.ML

增强平衡权重作为线性回归

Augmented balancing weights as linear regression

  • UC Berkeley(加州大学伯克利分校)
  • Ghent Univ.(根特大学)
  • Johns Hopkins Univ.(约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

David Bruns-Smith, Oliver Dukes, Avi Feller, Elizabeth L. Ogburn

更新

AI总结:

本文证明增强平衡权重估计器在线性基条件下等价于单一线性回归,并推广到岭回归和 lasso 情形,揭示其退化、欠平滑及双重选择性质,为理解该类估计器性能提供新框架。

AI中文摘要:

我们提供了增强平衡权重(也称为自动去偏机器学习,AutoDML)的一种新颖刻画。这些流行的双重稳健或去偏机器学习估计器将结果建模与平衡权重相结合——这些权重直接实现协变量平衡,而无需估计和求逆倾向得分。当结果模型和权重模型都在某个(可能无限的)基上呈线性时,我们证明增强估计器等价于一个单一线性模型,其系数结合了原始结果模型的系数和在同一数据上进行无惩罚普通最小二乘(OLS)拟合得到的系数。我们发现,在正则化参数的某些选择下,增强估计器常常退化为仅由 OLS 估计器构成;例如在对 Lalonde 1986 数据集的重新分析中就出现了这种情况。随后,我们将这些结果扩展到结果模型和权重模型的特定选择。我们首先证明,对结果模型和权重模型都使用(核)岭回归的增强估计器等价于一个单一的、欠平滑的(核)岭回归。这在有限样本中数值上成立,并为欠平滑和渐近收敛速率的新分析奠定了基础。当权重模型改为 lasso 惩罚回归时,我们给出了特殊情形下的闭式表达式,并展示了一种“双重选择”性质。我们的框架打开了这一日益流行的估计器类别的黑箱,弥合了关于欠平滑估计器和双重稳健估计器的半参数效率的现有结果之间的差距,并为增强平衡权重的性能提供了新的见解。

英文摘要:

We provide a novel characterization of augmented balancing weights, also known as automatic debiased machine learning (AutoDML). These popular doubly robust or de-biased machine learning estimators combine outcome modeling with balancing weights - weights that achieve covariate balance directly in lieu of estimating and inverting the propensity score. When the outcome and weighting models are both linear in some (possibly infinite) basis, we show that the augmented estimator is equivalent to a single linear model with coefficients that combine the coefficients from the original outcome model and coefficients from an unpenalized ordinary least squares (OLS) fit on the same data. We see that, under certain choices of regularization parameters, the augmented estimator often collapses to the OLS estimator alone; this occurs for example in a re-analysis of the Lalonde 1986 dataset. We then extend these results to specific choices of outcome and weighting models. We first show that the augmented estimator that uses (kernel) ridge regression for both outcome and weighting models is equivalent to a single, undersmoothed (kernel) ridge regression. This holds numerically in finite samples and lays the groundwork for a novel analysis of undersmoothing and asymptotic rates of convergence. When the weighting model is instead lasso-penalized regression, we give closed-form expressions for special cases and demonstrate a ``double selection'' property. Our framework opens the black box on this increasingly popular class of estimators, bridges the gap between existing results on the semiparametric efficiency of undersmoothed and doubly robust estimators, and provides new insights into the performance of augmented balancing weights.

↑