AI 中文总结
针对具有二元响应的高维广义线性模型变量选择问题,提出BMT方法,该方法考虑多重检验,能以高概率选到真实系数非零的协变量,具有神谕性质,实验表明其优于其他方法,还给出了有良好样本外表现的通胀预测模型。
AI 中文摘要
本文提出了一种具有多重检验的非线性提升(BMT)方法,用于具有二元响应的高维广义线性模型中的变量选择。在BMT过程的每个阶段,基于先前阶段已选的协变量,仅添加最显著的协变量来更新模型,同时考虑问题的多重检验性质。研究表明,在规定条件下,BMT过程以趋于1的概率选择所有真实系数非零的协变量,且不选其他协变量。此外,该过程具有神谕性质,即BMT后的模型参数最大似然估计渐近等同于预先知道正确稀疏模型的神谕估计。蒙特卡罗实验表明BMT优于其他竞争方法,具有高协变量选择准确性和低参数估计误差。一个实证例子说明BMT为美国通胀在12个月内超过给定阈值的概率提供了一个预测模型,具有很好的样本外表现。
英文摘要
This paper proposes a nonlinear boosting with multiple testing (BMT) approach to variable selection in high-dimensional generalised linear models with binary responses. At each stage of the BMT procedure, the model is updated by adding only the most significant covariate, conditional on those already selected in previous stages, while taking into account the multiple testing nature of the problem. It is shown that, under the stated conditions, the BMT procedure selects all covariates whose true coefficients are nonzero, and no other covariates, with probability tending to one. Furthermore, the procedure enjoys an oracle property, in the sense that the post-BMT maximum likelihood estimator of the parameters of the model is asymptotically equivalent to an oracle estimator that knows the correct sparse model in advance. Monte Carlo experiments demonstrate that BMT outperforms competing methods, delivering high covariate-selection accuracy and low parameter estimation error. An empirical example illustrates that BMT delivers a predictive model for the probability that U.S. inflation exceeds a given threshold over a 12-month horizon which has very good out-of-sample performance.