DeepGOF-1: 一种用于逻辑回归的预训练卷积拟合优度检验,具有可计算的相容性证书
[DeepGOF] Where Does a Logistic Risk Model Fail? An Audited Neural Goodness-of-Fit Test for Model Development and External Validation
浏览论文内容
中文总结 AI 辅助
针对逻辑回归拟合优度检验在小样本下水平漂移的问题,提出基于预训练卷积网络的检验方法,利用自助法校准保证水平精确性,并提供可计算的一致性证书,在基准上表现稳定且功效优越。
中文摘要 AI 辅助
逻辑回归的拟合优度检验在最需要的地方最不可靠:在小样本下,其水平偏离名义水平,而组合检验会加剧这种偏离。我们提出了一种检验方法,其统计量是一个卷积网络,该网络在模拟偏离上训练一次,将失拟视为图像:一个在协变量秩上的标准化残差网格。分析师从不训练。网络以冻结状态发布,p值是观察得分在分析师自身自助法中的秩,因此水平是校准的性质,而不是网络学习到的内容。我们在枢轴性下证明了精确性,在没有枢轴性时证明了渐近精确性,并证明了一个一致性定理,其关键条件可以从冻结权重在一次前向传播中计算出来,从而为每个备择假设提供证书;我们还测量了检验的盲锥。在一个预先声明的六十单元网格上,部署的水平在五十八个单元中保持在名义带内,而该网格的功效标准未通过;在一个已发表的基准上,它在不同设置和样本量下的十三个水平中是最稳定的,并且该检验在每个样本量下都优于所有分区检验,同时总体排名第七。一个银行失败应用和一个版本化发布结束了本文。
英文摘要
Goodness-of-fit tests for logistic regression are routine in clinical risk modelling, yet the classical tests say whether a model misfits, not where. A pretrained network is not a test until its level, power and blind spots are established. We audit DeepGOF-1, which renders the residuals of a fitted logistic model as a map over the ranks of two covariates, scores the map with a frozen convolutional network, and calibrates the score by the analyst's parametric bootstrap. The level is first-order valid under standard bootstrap regularity and two conditions, consistency against a named alternative is decided by one forward pass, and local power has an explicit limit. In simulation the level holds from 20 patients to 8,873; the network adds power over a chi-square statistic of the same map when misfit is sparse among many covariates; the map finds a missed interaction that one-covariate diagnostics cannot see; and, applied to a published risk model on new patients, it is an exactly valid test of calibration within patient subgroups, where the calibration belt has little power. In the SUPPORT study of 8,873 inpatients, DeepGOF-1 rejects a linear in-hospital mortality model, its map shows the U-shaped risks of blood pressure and respiratory rate that a calibration curve hides, and it guides the repairs. It ran in 13 seconds; the projection test, the strongest omnibus rival, took 9 hours 53 minutes and 4.5 GB on the same data, too slow for routine use. The test is deepgof1() in the R package ebrahim.gof.
发表机构
- Alexandria University(亚历山大大学)
机构由 AI 辅助整理,请以论文原文为准。