arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EDGE:概率二分类器校准的闭式定向检验

[EDGE] A grouped calibration test for logistic regression that tolerates a few corrupted records

Ebrahim Khaled Ebrahim, Osama Abd El-Aziz Hussein, Ahmed El-Kotory

arXiv 2608.20511首次发表:更新:

发表机构

Department of Applied Statistics, Alexandria University(亚历山大大学应用统计系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对概率二分类器校准的问题,提出闭式定向检验EDGE,利用分箱表与残差投影构建原分布,在多数场景下表现优于同类分箱检验,且对稀疏性鲁棒。

AI 中文摘要

概率二分类器几乎处处由判别性——准确率、ROC曲线及其曲线下面积来评判。每一个这类准则都对预测概率的单调失真具有不变性,因此分类器即便能完美排序,返回的概率仍可能严重错误。校准是决策所需的属性,该领域用于校准的工具是带可靠性图的分箱期望校准误差,它是描述性的:没有原分布,因此无法判断其显示的校准误差是真实存在还是噪声,且依赖于分箱。我们提出EDGE,这是针对经典概率分类器——逻辑回归的校准检验。EDGE读取与可靠性图绘制所用相同的分箱预测-观测表,并将其标准化分箱残差投影到预先指定的光滑校准失真形状的小子空间上。它的原分布是卡方变量的加权和,呈闭式形式,仅需遍历一次数据和一次小型特征分解:无需重新拟合、无需重采样、无需调优,因此可在交叉验证循环内运行。分箱也使其对连续特征产生的稀疏性具有鲁棒性。在链接和特征误设定的情况下,预先指定的默认版本在22种可检测场景中的19种里,在拟合指标上优于或与所有分箱检验持平,且在基于重新拟合的Stukel得分检验因20%至28%的稀疏样本而失效的情况下仍可计算。其诚实极限是粗糙的高频校准误差,这是综合统计量占优的极限,一个基础的分辨率论证表明,包括校准误差在内的所有分箱工具都具有该极限。

英文摘要

The Hosmer-Lemeshow calibration test groups patients by predicted risk for a valid reference and loses power that more groups cannot recover. Later tests weight each record's residual alone, which makes them fragile: on an otherwise correct model, one record in a thousand with a corrupted covariate raises the false-alarm rate of Stukel's score test from 5% to 12%, and ten to 63%. Pooling on equal-size groups bounds a wrong record's influence, dividing its residual by its groupmates' variance; without pooling, the same directions are as fragile as Stukel's test after a refit. EDGE, the test proposed here, projects the grouped residuals onto a small basis of calibration shapes. Its degrees of freedom do not depend on the number of groups, so the partition becomes a setting, fine-grained for power and coarse when records may be wrong, and its reference needs no resampling. In a prespecified simulation EDGE was comparable in power to Stukel's test on clean data and 0.064 more powerful on average than the ten-group Hosmer-Lemeshow test. With one and ten exaggerated covariates in a thousand its false-alarm rates were 0.056 and 0.097, and with ten groups below 0.10 up to about ten exaggerated or five reversed. In external validation, at ten groups, its average power was that of Cox's recalibration test and the GiViTI belt, and its false-alarm rate stayed below 0.10 with ten reversed records in a thousand, where theirs did not. In a 70,000-patient cohort its verdict did not change with the partition.

Comments25 pages, 6 figures, 6 tables; Supporting Information as an ancillary file. v2: revised and retitled, with new studies of corrupted records and external validation and a clinical cohort. R package ebrahim.gof (CRAN). Archive: https://doi.org/10.5281/zenodo.23079225. Develops a method introduced in one chapter of the first author's M.Sc. thesis (arXiv:2608.11140)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑