arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28608cs.LGq-bio.QM

KAISEN:临床风险模型可复现的子群体公平性审计

KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

Sparsh Roy, Samuel Girmachew, Nishita Chavan

首次发表
浏览论文内容

中文总结 AI 辅助

KAISEN是一个五阶段临床风险模型子群体公平性审计流程,经合成基准测试揭示了审计各环节的特性与局限性,相关复现资源已公开。

中文摘要 AI 辅助

临床风险模型通常在整体上表现出良好性能,但在不同患者子群体中会产生显著不同的错误率。已有研究提出了审计流程来检测这种情况,但这些流程的组成部分很少经过压力测试,因此尚不清楚审计的哪些部分可以信任,以及在什么条件下可以信任。我们提出了KAISEN,这是一个五阶段的审计流程,涵盖子群体分层、差异测量、机制诊断、事后缓解和漂移监测,我们在包含16个疾病任务、来自《健康人民2030》(Healthy People 2030)的15个社会决定因素轴以及3个预先指定的交集的合成基准上对其进行了失效测试。得出四个发现:(i)显著性跟踪每个轴与其自身最小可检测效应的差距:15个轴的显著性数量与原始均衡赔率差异(EOD)之间的秩相关系数为ρ=0.56,当EOD按该下限标准化后,该系数上升至ρ=0.78。(ii)每组阈值优化在48次保留运行中均降低了EOD(配对差值=-0.285,95%置信区间[-0.313, -0.252]),而组内Platt缩放作为更好的校准器,对EOD的表现类似抛硬币(48次运行中有19次得到改善,95%置信区间[0.26, 0.55]),平均效应接近零,因此审计应报告的是方差而非平均值。(iii)机制诊断正确分类了144个受控案例,但在代理误设情况下未检测到48个模型驱动案例中的任何一个,且没有信号表明其失效。(iv)CUSUM失效和虚警更多地与队列实现而非疾病相关:在参考阈值下,所有27次虚警和8次漏检变化中的7次来自不同的随机种子(卡方检验p=0.002),因此在一个队列上调整的阈值无法迁移。所有结果均基于具有已知真实值的合成数据,不构成临床有效性证明。我们发布了可复现所有数值的代码、人工制品和脚本。

英文摘要

Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections. Four findings follow. (i) Significance tracks each axis's gap against its own minimum detectable effect: rank correlation between significance count and raw equalized-odds difference (EOD) across the 15 axes is rho = 0.56, rising to rho = 0.78 once EOD is standardized by that floor. (ii) Per-group threshold optimization reduces EOD in 48 of 48 held-out runs (paired delta = -0.285, 95% CI [-0.313, -0.252]), while group-wise Platt scaling -- the better calibrator -- behaves as a coin flip on EOD (19 of 48 runs improved, 95% CI [0.26, 0.55]) with mean effect near zero, so what an audit should report is the variance, not the average. (iii) The mechanism diagnostic classifies 144 of 144 controlled cases correctly but recovers none of 48 model-driven cases under proxy misspecification, with no signal that it failed. (iv) CUSUM failures and false alarms track cohort realization far more than disease: at the reference threshold, all 27 false alarms and 7 of 8 missed shifts come from different seeds (chi-squared p = 0.002), so a threshold tuned on one cohort fails to transfer. All results are synthetic with known ground truth and do not establish clinical validity. Code, artifacts, and scripts reproducing every number are released.

发表机构

  • Massachusetts Institute of Technology(麻省理工学院)
  • Hopewell Valley Central High School(霍普韦尔谷中央高中)
  • East Brunswick High School(东布伦瑞克高中)

机构由 AI 辅助整理,请以论文原文为准。

↑