逻辑回归分类器的拟合优度和校准算法基准测试:稀疏数据下的大规模模拟研究
Benchmarking Goodness-of-Fit and Calibration Algorithms for Logistic Regression Classifiers: A Large-Scale Simulation Study under Sparse Data
浏览论文内容
中文总结 AI 辅助
研究逻辑回归分类器拟合优度和校准算法,通过开源R包实现二十多种测试,在多种场景下评估,发现紧凑核心测试平衡最佳且比Hosmer-Lemeshow测试更强大,为评估逻辑回归拟合提供实用指南。
中文摘要 AI 辅助
二元逻辑回归是最广泛使用的分类算法之一,但其预测概率需校准才可信。经典检验在现代连续预测变量和稀疏数据设置中失效。四十年来有众多替代算法,本文提供统一分类法和大规模可重复模拟基准,在开源R包中实现二十多种测试。在五种协变量分布和四种错误设定场景下评估,测量第一类错误和功效。一些经典测试宽松,另一些功效低。紧凑核心测试在正确规模和高功效间平衡最佳,比常用的Hosmer-Lemeshow测试更强大。低体重应用强化了这一点,未考虑交互作用的模型几乎能通过所有测试,只有结合敏感测试和校准曲线才能暴露。本文将结果转化为评估逻辑回归拟合的实用循证指南。
英文摘要
Binary logistic regression is among the most widely used classification algorithms, yet a classifier is only trustworthy if its predicted probabilities are well calibrated. The classical checks -- the Pearson chi-square and deviance statistics -- break down precisely in the modern setting where predictors are continuous and the data are sparse (one covariate pattern per observation). Four decades of research have produced dozens of alternative goodness-of-fit and calibration algorithms, yet practitioners still default to the Hosmer-Lemeshow test because it ships with their software. This paper provides a unified taxonomy and a large-scale, reproducible simulation benchmark; more than twenty tests are implemented in the open-source R package ebrahim.gof. We evaluate them across five covariate distributions and four misspecification scenarios, with 10,000 replications each, measuring both Type I error and power. Several classical tests prove liberal, rejecting correct models far too often, while others have little power. A compact core -- McCullagh, Osius-Rojek, le Cessie-van Houwelingen, Stute-Zhu, and the GiViTI calibration test -- delivers the best balance of correct size and high power, and is consistently more powerful than the ubiquitous Hosmer-Lemeshow test. A low-birth-weight application reinforces the point: a model with omitted interactions slips past nearly every test, exposed only by pairing sensitive tests with a calibration (reliability) curve. We translate these findings into practical, evidence-based guidance for assessing logistic regression fit.