稀疏数据下逻辑回归的拟合优度检验与校准机器学习算法
Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data
AI总结:
本论文针对稀疏数据下的逻辑回归,比较约30种统计检验与机器学习校准算法,发现多种检验需结合可视化检查,才能有效评估模型拟合优度。
AI中文摘要:
在将逻辑回归模型用于推断之前,评估其拟合优度是一项关键前提。然而,当数据“稀疏”时,卡方检验和偏差检验等拟合优度(GOF)检验往往会给出无效结果;稀疏数据是年龄或体重等连续预测变量的常见问题,此时渐近分布假设无法满足。本论文研究了分组数据和稀疏数据下二元逻辑回归的经典GOF检验,比较了约30种统计检验与机器学习校准算法,涵盖经典卡方检验、Hosmer-Lemeshow变体、标准化皮尔逊统计量、协变量空间划分、基于平滑的方法,以及当代校准机器学习和自助法程序。在固定样本量下,GiViTI校准检验(2016)、McCullagh(1989)、Osius-Rojek(1992)、le Cessie(1995)和Stute-Zhu(2002)经实证证明具有较强功效,在正确识别不良模型(高实证功效)与不对良好模型发出误报(正确的实证I类错误)之间实现了平衡。仅依赖形式化方法是不够的:校准图等可视化诊断是检测形式化检验常忽略的模型缺陷的重要探索步骤。对真实数据(低出生体重数据集)的应用表明,当这些检验面临实际数据集的复杂性时,许多检验无法给出有效结论。主要结论是,模型评估需要结合多种强大的统计检验,并辅以对模型校准的仔细可视化检查。
英文摘要:
Assessing the goodness-of-fit of a logistic regression model is a critical prerequisite before the model is used for inference. However, goodness-of-fit (GOF) tests such as the chi-square and deviance tests often give invalid results when the data are "sparse" -- a common issue with continuous predictors like age or weight, where the asymptotic distributional assumptions are not satisfied. This thesis studies classical GOF tests for binary logistic regression under both grouped and sparse data, comparing about 30 statistical tests and machine-learning calibration algorithms. These span the classical chi-square and Hosmer-Lemeshow variants, standardized Pearson statistics, covariate-space partitioning, smoothing-based methods, and contemporary calibration machine-learning and bootstrap procedures. At a fixed size, the GiViTI calibration test (2016), McCullagh (1989), Osius-Rojek (1992), le Cessie (1995) and Stute-Zhu (2002) proved empirically powerful, balancing correct identification of bad models (high empirical power) against not raising false alarms on good models (correct empirical Type I error). Relying on formal methods alone is insufficient: visual diagnostics such as calibration plots are a vital exploratory step for detecting model deficiencies that formal tests often overlook. An application to real data (the Low Birth Weight dataset) shows that many of these tests fail to give valid conclusions when exposed to the complexities of actual datasets. The main conclusion is that model assessment requires a combination of several powerful statistical tests alongside careful visual inspection of model calibration.