AI 中文总结
该研究针对高维Lasso推断,提出通过标准化数据集的方法,在更弱假设下实现去偏Lasso类方法的有效推断,降低了对数据尾分布等强假设的要求。
AI 中文摘要
我们研究高维Lasso估计量的推断问题,已有多种方法(包括“双重选择”技术及多个版本的“去偏Lasso”)被提出并取得显著成效。然而,大多数理论保证都对潜在数据过程和线性回归模型的误差施加了强假设,例如亚高斯设计、误差与数据本身相互独立。我们证明,对数据集进行“标准化”(这是惩罚回归实践中的自然步骤),可在弱得多的假设下得到相同结果,且仅因未假设轻尾分布付出微小代价。该步骤带来的关键技术点是利用自正则化过程的集中性质。重要的是,我们针对与“去偏Lasso”密切相关的两种不同方法证明了结果,其中第二种方法在错误设定的线性模型下,即使在类似“双重选择”文献的温和稀疏条件下,也能进行有效推断。
英文摘要
We consider the problem of high-dimensional inference with the lasso estimator. Different methods including 'double selection' techniques and multiple versions of the 'debiased lasso' have been proposed for this task with noticeable success. However, most guarantees assume strong hypotheses on the underlying data process and the errors in the linear regression model, such as subgaussian designs and independence between errors and the data itself. We show that 'standardizing' one's dataset -- a natural procedure in practical penalized regression -- leads to the same results under much weaker hypotheses, paying only a small price for not assuming light tails. The key technical point allowed by this step is exploiting the concentration properties of self-normalized processes. Importantly, we prove our results for two different methods closely related to the 'debiased lasso'. The second method performs valid inference even for a misspecified linear model, under mild sparsity conditions similar to the 'double selection' literature.