arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06413stat.MEstat.APstat.CO

收缩使Hosmer-Lemeshow检验失效:惩罚逻辑回归的拟合优度检验及其在青光眼诊断中的应用

Shrinkage invalidates the Hosmer-Lemeshow test: goodness of fit for penalized logistic regression, with an application to glaucoma diagnosis

Ebrahim Khaled Ebrahim

首次发表
浏览论文内容

中文总结 AI 辅助

惩罚逻辑回归下Hosmer-Lemeshow检验因收缩失效,提出收缩校正版本,并在青光眼诊断中显示模型需重新校准而非重建。

中文摘要 AI 辅助

临床预测模型越来越多地采用惩罚逻辑回归进行拟合,因为共线性或候选预测因子过多会使最大似然估计不稳定甚至无法实现。此时,校准几乎总是通过分组拟合优度检验(如Hosmer-Lemeshow检验)来评估。我们证明这种组合是无效的。在岭回归下,分组标准化残差因收缩而产生非中心性,因此实践中使用的参考分布是错误的;在使拟合概率改善最大的惩罚值下,该检验在92%至100%的情况下拒绝正确设定的模型。我们推导了修正的分布规律,并定义了收缩校正的Hosmer-Lemeshow检验,该检验减去非中心性的估计值,从而在一阶意义上精确恢复最大似然参考分布,并通过预枢轴化(prepivoting)使其有效,我们测量了由此带来的功效代价。我们还给出了衰减规律,该规律决定了在线性预测器必须被估计时,任何此类检验能够检测到什么。在通过共焦激光断层扫描诊断青光眼时,最大似然估计不存在,校正后的检验发现失拟合的证据强度减弱了三个数量级以上:拟合的风险过于平坦而非排序错误,因此模型需要重新校准而非重建。

英文摘要

Clinical prediction models are increasingly fitted by penalized logistic regression, because collinearity or many candidate predictors makes maximum likelihood unstable or impossible. Calibration is then almost always assessed by a grouped goodness-of-fit test such as the Hosmer-Lemeshow test. We show that this combination is invalid. Under ridge regression the grouped standardized residuals acquire a non-centrality induced by shrinkage, so the reference distribution used in practice is wrong, and at the penalty that most improves the fitted probabilities the test rejects correctly specified models between 92 and 100 per cent of the time. We derive the corrected law and define the shrinkage-corrected Hosmer-Lemeshow test, which subtracts an estimate of that non-centrality, restoring the maximum likelihood reference exactly to first order, and is made valid by prepivoting at a power cost we measure. We also give the attenuation law governing what any such test can detect once the linear predictor must be estimated. In glaucoma diagnosis by confocal laser tomography, where the maximum likelihood estimate does not exist, the corrected test finds the evidence for misfit weaker by more than three orders of magnitude: the fitted risks are too flat rather than mis-ordered, so the model needs recalibration rather than rebuilding.

补充信息

↑