发表机构
University of Rochester; Texas A&M University; University of California, Los Angeles(罗切斯特大学; 德克萨斯农工大学; 加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文推导IV-LASSO估计量的近似均方误差,揭示偏差-方差权衡,提出基于AMSE的惩罚选择方法,在模拟中比交叉验证降低最多三分之一均方误差。
AI 中文摘要
使用最小绝对收缩和选择算子(LASSO)估计工具变量(IV)模型的第一阶段,需要选择技术工具字典和惩罚水平。一阶渐近理论对这些选择不提供指导,因为任何一致的实施都会产生具有相同极限分布的结构参数估计量。然而在有限样本中,这些选择可能对最终的结构参数估计产生实质性影响。在具有单一内生回归量和同方差高斯误差的模型中,我们利用一阶和二阶Stein恒等式推导出工具变量LASSO(IV-LASSO)估计量的近似均方误差(AMSE),该误差可被一致估计并用于对预指定的字典-惩罚候选列表进行排序。AMSE揭示了偏差-方差权衡:更复杂的第一阶段拟合能更好地逼近内生变量的条件均值,但也与结构误差更相关,其中复杂度由LASSO拟合的自由度衡量。该偏差的权重随回归量的内生性增加而上升,而插入法和交叉验证的惩罚规则均未考虑这一数量。尽管AMSE是在高斯模型中推导的,但通过最小化可行AMSE准则进行惩罚选择,在与Gilchrist和Sands(2016)数据校准的高斯和非高斯模拟设计中,相比交叉验证和插入法惩罚规则,均方误差最多可降低三分之一。
英文摘要
Estimating the first stage of an instrumental variables (IV) model with the least absolute shrinkage and selection operator (LASSO) requires choosing a dictionary of technical instruments and a penalty level. First-order asymptotic theory offers no guidance on these choices, as any consistent implementation yields a structural parameter estimator with the same limiting distribution. In finite samples, however, these choices can have a substantial impact on the resulting structural parameter estimate. Working in a model with a single endogenous regressor and homoskedastic Gaussian errors, we use first- and second-order Stein identities to derive the approximate mean squared error (AMSE) of the instrumental-variables LASSO (IV-LASSO) estimator, which can be consistently estimated and used to rank a prespecified list of dictionary-penalty candidates. The AMSE reveals a bias-variance trade-off: more complex first-stage fits better approximate the conditional mean of the endogenous variable but are also more correlated with the structural errors, with complexity measured by the degrees of freedom of the LASSO fit. The weight on this bias rises with the endogeneity of the regressor, a quantity that neither plug-in nor cross-validation penalty rules take into account. Despite the AMSE being derived in a Gaussian model, penalty selection by minimizing the feasible AMSE criterion delivers up to a one-third lower mean squared error compared to cross-validation and plug-in penalty rules in Gaussian and non-Gaussian simulation designs calibrated to the data of Gilchrist and Sands (2016).