arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30174stat.ME

广义线性模型中变量选择的平滑信息准则

Smooth Information Criterion for Variable Selection in Generalised Linear Models

Andrew McInerney

首次发表
浏览论文内容

中文总结 AI 辅助

针对广义线性模型变量选择,提出平滑信息准则(SIC),以可微近似替代离散模型维度项,在模拟和真实数据中紧密复现穷举BIC选择,计算高效且性能有竞争力。

中文摘要 AI 辅助

使用信息准则进行变量选择具有明确的统计目标,但需要对候选模型进行离散搜索。平滑信息准则(SIC)用可微近似替代不连续的模型维度项,并采用ε-望远镜连续化策略逐步锐化该近似,无需数据驱动地选择正则化强度参数。我们将SIC发展为广义线性模型(GLMs)中系数级变量选择的通用程序,并用以解决一个核心问题:平滑优化在多大程度上忠实再现相应的离散信息准则选择问题?聚焦于BIC,我们在可行情况下将SIC直接与穷举子集选择进行基准比较,使用精确支持一致性、BIC差异以及跨不同信号强度边界的选择行为。在高斯、二项和泊松回归的模拟中,SIC紧密再现穷举BIC选择,并跟踪精确的BIC选择边界。相对于逐步BIC、LASSO、SCAD和MCP,SIC在保持稀疏模型的同时产生有竞争力的变量选择性能,预测性能在各方法间大致相当。相对于逐步BIC的计算优势随预测变量维度增加而增大。在具有16个候选预测变量的真实数据应用中,SIC以穷举枚举计算成本的一小部分,在所有65,536个支持中恢复全局BIC最优模型。

英文摘要

Variable selection using information criteria has an explicit statistical target but requires discrete search over candidate models. The smooth information criterion (SIC) replaces the discontinuous model-dimension term by a differentiable approximation, with an $ε$-telescoping continuation strategy progressively sharpening this approximation without data-driven selection of a regularisation-strength parameter. We develop SIC as a general procedure for coefficient-level variable selection in generalised linear models (GLMs) and use it to address a central question: how faithfully does smooth optimisation reproduce the corresponding discrete information-criterion selection problem? Focusing on BIC, we benchmark SIC directly against exhaustive subset selection where feasible, using exact support agreement, BIC difference and selection behaviour across a varying signal-strength boundary. Simulations in Gaussian, binomial and Poisson regression show that SIC closely reproduces exhaustive BIC selection and tracks the exact BIC selection boundary. Relative to stepwise BIC, LASSO, SCAD and MCP, SIC produces competitive variable-selection performance while retaining sparse models, with predictive performance broadly comparable across methods. Computational advantages over stepwise BIC increase with predictor dimension. In a real-data application with 16 candidate predictors, SIC recovers the globally BIC-optimal model among all 65,536 supports at a small fraction of the computational cost of exhaustive enumeration.

发表机构

  • School of Medicine, University of Limerick(利默里克大学医学院)

机构由 AI 辅助整理,请以论文原文为准。

↑