发表机构
Chiang Mai University(清迈大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究基于GUSTO-1数据比较逻辑回归与人工神经网络,发现判别能力指标比个体预测更早稳定,提示模型开发者需评估预测层面的稳定性。
AI 中文摘要
模型验证通常用于量化和校正性能估计中的过拟合,但它并未完全解决由略微不同的训练样本所开发模型的稳定性问题。我们研究了稳定的判别能力是否意味着稳定的个体预测,并比较了逻辑回归(LR)与人工神经网络(ANN)。利用GUSTO-I试验中40,830名参与者的数据,以30天死亡率为结局,我们创建了六个样本量场景,其每个变量事件数(EPV)各不相同。在每个场景中,使用八个预测变量开发LR和ANN模型,并采用200次bootstrap重采样来评估判别能力、平均绝对预测误差(MAPE)和分类不稳定性指数(CII)的稳定性。对于LR,判别能力在EPV为44.50时趋于稳定,而MAPE和CII直到EPV为89.12时才稳定。ANN的判别能力波动更大,其乐观性仅在EPV为178.25时才稳定,且ANN预测在EPV为4.50时严重校准不良。总体而言,两种方法的判别能力均比个体预测更早稳定,但ANN表现出更大的不稳定性和更高的样本量需求。因此,模型开发者应评估并报告预测层面的稳定性,因为仅凭判别指标可能无法反映个体预测的可靠性。
英文摘要
Model validation is routinely used to quantify and correct for overfitting in performance estimates, but it does not fully address the stability of models developed from slightly different training samples. We investigated whether stable discrimination implies stable individual predictions and compared logistic regression (LR) with artificial neural networks (ANNs). Using data from 40,830 participants in the GUSTO-I trial, with 30-day mortality as the outcome, we created six sample-size scenarios with varying events per variable (EPV). In each scenario, LR and ANN models were developed using eight predictors, and 200 bootstrap resamples were used to assess stability in discrimination, mean absolute prediction error (MAPE), and the classification instability index (CII). For LR, discrimination stabilized at an EPV of 44.50, whereas MAPE and CII did not stabilize until an EPV of 89.12. ANN discrimination was more volatile, with optimism stabilizing only at an EPV of 178.25, and ANN predictions were severely miscalibrated at an EPV of 4.50. Overall, discrimination stabilized earlier than individual predictions for both approaches, with greater instability and larger sample-size requirements for ANN. Model developers should therefore assess and report prediction-level stability, because discrimination metrics alone may not reflect the reliability of individual predictions.
Comments21 pages