arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

识别临床AI中预测多重性亚组脆弱性的统计框架

A statistical framework for identifying subgroup vulnerability to predictive multiplicity in clinical AI

Enock Adu Bonsu

arXiv 2609.37064首次发表:更新:

发表机构

University of Arizona(亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出统计框架V(S),结合模型不一致下界与临床严重性,在MIMIC-IV和eICU-CRD队列中审计158个亚组,识别心脏病亚组脆弱性,提供可复现的亚组排序方法。

AI 中文摘要

基于相同数据训练的AI模型可能对患者风险产生不一致的预测,且这种不一致可能集中在临床重要的亚组中。我们提出了V(S),一个基于统计的脆弱性指数,将模型不一致的可观察下界证据与临床严重性相结合,并开发了用于审计预设亚组的推断和多重性调整程序。我们将该框架应用于两个大型重症监护队列:MIMIC-IV(n = 65,078)用于模型开发,eICU-CRD(n = 188,230次入院,208家医院)用于外部验证,在158个预设亚组中比较了随机森林与逻辑回归。两个主要模型并未都满足预设的epsilon = 0.02的Rashomon集容差:逻辑回归的AUC比最佳候选模型AUC低0.0488。因此,RF-LR判别差距被解释为两个特定模型之间的不一致,而非整个Rashomon集的保证下界。九个亚组的判别差距与预设的临床下限可区分。年龄≥80岁且合并心脏病的亚组具有最大的V(S)点估计值(0.307),但统计功效不足,未满足完整的高优先级决策规则。单变量心脏病亚组(V(S) = 0.193,95% CI [0.163, 0.223])是唯一具有足够功效且统计上可区分的亚组。事后分析确定乳酸对两个模型均重要,但未建立不一致的因果解释。四项模拟研究量化了所提程序的操作特征,包括小样本检测率膨胀和Wald区间覆盖率不完美。该框架提供了一种可复现的方法来对亚组对模型不一致的脆弱性进行排序,同时将探索性信号与充分支持的发现区分开来。

英文摘要

AI models trained on the same data can disagree about patient risk, with disagreement potentially concentrated in clinically important subgroups. We propose V(S), a statistically grounded vulnerability index combining an observable lower-bound witness of model disagreement with clinical severity, and develop inference and multiplicity-adjustment procedures for auditing prespecified subgroups. We applied the framework to two large critical-care cohorts, MIMIC-IV (n = 65,078) for model development and eICU-CRD (n = 188,230 admissions, 208 hospitals) for external validation, comparing a random forest with logistic regression across 158 prespecified subgroups. The two primary models did not both satisfy the prespecified epsilon = 0.02 Rashomon-set tolerance: the logistic-regression AUC was 0.0488 below the best candidate-model AUC. The RF-LR discrimination gap is thus interpreted as disagreement between two specific models, not as a guaranteed lower bound on the full Rashomon set. Nine subgroups had discrimination gaps distinguishable from a prespecified clinical floor. The age >=80 and cardiac subgroup had the largest point estimate of V(S) (0.307), but was underpowered and did not meet the full high-priority decision rule. The univariate cardiac subgroup (V(S) = 0.193, 95% CI [0.163, 0.223]) was the only statistically distinguishable subgroup with adequate power. Post hoc analyses identified lactate as important for both models but did not establish a causal explanation for the disagreement. Four simulation studies quantified operating characteristics of the proposed procedures, including inflated small-sample detection rates and imperfect Wald-interval coverage. The framework offers a reproducible approach for ranking subgroup vulnerability to model disagreement while separating exploratory signals from adequately supported findings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑