临床预测AI中性能指标的可折叠性
Collapsibility of Performance Metrics in Clinical Predictive AI
- University of Oxford(牛津大学)
- KU Leuven(鲁汶大学)
- University of Birmingham(伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究分析临床预测AI的15种性能指标的可折叠性,发现AUC等5种指标非可折叠、10种可折叠,指出非可折叠性会误导公平性评估,明确报告其可折叠性可提升评估的可解释性与透明度。
AI中文摘要:
背景:对预测人工智能(AI)的总体水平评估可能会掩盖各亚组之间的性能差异,公平性评估通常依赖于各亚组的性能分析。然而,部分性能指标具有非可折叠性,即总体人群的性能值不等于各亚组特定值的加权平均值。目的:研究预测AI中常用性能指标的可折叠性,重点关注受试者工作特征曲线下面积(AUC,又称c统计量)。方法:我们通过将15种性能指标表达为其分层特定值的线性组合,或对非可折叠指标提供受辛普森悖论启发的反例作为形式化反证,来研究这些指标的可折叠性。结果:5种性能指标(AUC、校准截距、校准斜率、预期校准误差和Nagelkerke R²)被证明具有非可折叠性,另外10种指标(观测期望比(O:E ratio)、对数损失(logloss)、布里尔分数(Brier score)、准确率(accuracy)、F1分数(F1-score)、真阳性率(true positive rate)、真阴性率(true negative rate)、阳性预测值(positive predictive value)、阴性预测值(negative predictive value)和净获益(net benefit))被证明具有可折叠性。AUC具有非可折叠性,因为当存在亚人群时,它会分解为组内和组间的AUC项,导致其总体值可能落在各亚组特定AUC值的范围之外。结论:性能指标的非可折叠性对报告、模型评估和公平性评估具有重要影响,它可能会在亚组性能与总体性能之间产生虚假差异,从而误导公平性评估。明确承认并报告性能指标的可折叠性可提高公平性评估的可解释性和透明度。
英文摘要:
Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness evaluations commonly rely on performance analyses across subgroups. However, some performance metrics are non-collapsible, meaning that the overall population performance value does not equal the weighted average of subgroup specific values. Objective: To examine the collapsibility properties of commonly reported performance metrics in predictive AI, with a focus on the area under the receiver operating characteristic curve (AUC, also known as c-statistic). Methods: We investigate the collapsibility of 15 performance metrics, either by expressing each metric as a linear combination of its stratum specific values or, where non-collapsible, by providing a counterexample inspired by Simpson's paradox as a formal disproof. Results: Five performance metrics (AUC, calibration intercept, calibration slope, expected calibration error, and Nagelkerke R^2) are shown to be non-collapsible, and ten (O:E ratio, logloss, Brier score, accuracy, F1-score, true positive rate, true negative rate, positive predictive value, negative predictive value, and net benefit) are shown to be collapsible. The AUC is shown to be non-collapsible because it decomposes into within- and cross-group AUC terms when subpopulations coexist, such that its overall value may fall outside the range of subgroup specific AUCs. Conclusions: Non-collapsibility of performance metrics has important consequences for reporting, model appraisal, and fairness evaluation. It can generate spurious differences between subgroup and overall performance, which may mislead fairness evaluations. Explicitly acknowledging and reporting the collapsibility properties of performance metrics improves both the interpretability and transparency of fairness assessments.