arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01333stat.MEstat.ML

用于数据驱动发现临床风险预测模型差异性能的统计框架

A Statistical Framework for Data-Driven Discovery of Differential Performance in Clinical Risk Prediction Models

Aidan Neher, Julian Wolfson

AI总结:

该研究提出utree统计框架,用于数据驱动识别临床风险预测模型的差异性能亚组,在模拟和真实心梗数据集上验证了其检测性能差异的有效性。

AI中文摘要:

采用人工智能(AI)和机器学习(ML)的预测模型正越来越多地用于医疗环境中的决策支持。这些模型在由种族、年龄、性别和其他因素定义的人群亚组中可能表现出差异性能,并导致不同的临床影响,因此近期对所谓“模型公平性”的研究十分密集。虽然已提出许多评估风险预测模型公平性的方法,但这些技术通常要求最终用户预先指定要评估公平性的组。然而在现实环境中,重要的模型性能差异可能出现在由多个交叉特征定义的未知亚组中。为解决该问题,我们提出了unfairness tree(utree),这是一种用于识别具有差异模型性能的亚组的数据驱动递归划分框架。在模拟中,utree表现出标称的经验I型错误率,且具备检测、量化和表征由高阶变量交互定义的性能差异的良好能力。在适配GUSTO-I急性心肌梗死试验数据集的六个死亡率风险模型中,utree识别出了亚组特定的性能模式,其中年龄、性别、血压和Killip分级始终与差异模型性能相关联。

英文摘要:

Predictive models employing artificial intelligence (AI) and machine learning (ML) are increasingly being used for decision support in healthcare settings. These models may exhibit differential performance across population subgroups defined by race, age, sex, and other factors and cause disparate clinical impacts, leading to intensive recent study of what has been termed "model fairness". While many methods have been proposed to assess risk prediction model fairness, these techniques generally require that the end user pre-specify the groups across which fairness is to be evaluated. In real-world settings, however, important model performance disparities may arise in unknown subgroups defined by multiple intersecting characteristics. To address this problem, we propose the unfairness tree (utree), a data-driven recursive partitioning framework for identifying subgroups with differential model performance. In simulations, the utree exhibits nominal empirical type I error rates and good ability to detect, quantify, and characterize performance discrepancies defined by higher-order variable interactions. In six mortality risk models fit to the GUSTO-I acute myocardial infarction trial dataset, utrees identified subgroup-specific performance patterns, with age, sex, blood pressure, and Killip class consistently associated with differential model performance.

↑