arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

跨医院分布偏移下考虑缺失感知的共形预测

Missingness-Aware Conformal Prediction Under Cross-Hospital Distribution Shift

Liang You, Dongwen Ou, Hengyu Shi, Siyuan Dai

arXiv 2609.30781首次发表:更新:

发表机构

University of Pittsburgh; Duke University; Xiamen University(匹兹堡大学; 杜克大学; 厦门大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对跨医院分布偏移下的死亡预测,提出缺失感知共形校准,按测量缺失分组进行蒙德里安校准,减少最差组覆盖差距,并揭示合并评估的局限性。

AI 中文摘要

临床测量仅对部分患者进行记录,且记录率在不同医院间存在差异,边缘共形覆盖并不能确保在由缺失性定义的组内实现覆盖。我们提出了一种在跨医院分布偏移下用于死亡预测的缺失感知共形校准程序。该程序在独立样本上选择一个测量指标,根据该指标是否被记录对患者进行分组,并在每组内应用蒙德里安校准,从而确保校准结果不被重复使用。我们在eICU数据集中的各医院之间以及MIMIC-IV中一家医院内的各护理单元之间,使用三种预测器对该程序进行了评估。相对于合并校准,该程序在所有六种设置下均减少了所选组上的平均最差组覆盖差距,中位数减少1.9个百分点;配对站点自助法区间在五种设置中排除零。这些收益并未均匀扩展。按预测风险进行校准在更广泛的缺失性组面板上实现了更小的差距,而当eICU医院被单独评估时,所有三种预测器的收益均缩小,其中一种预测器的收益符号发生反转。我们用医院层面的分解来解释这一差异。合并通过组份额与覆盖误差之间的协方差对医院进行重新加权,并允许符号相反的误差相互抵消:加权解释了反转现象,而抵消则解释了另外两种预测器大部分衰减的原因。构造的人群分布表明,即使在没有抽样噪声的情况下,合并评估和院内评估也可能对校准方法给出相反的排序。因此,仅凭合并改进不能确立医院内部的更好覆盖,即使校准组是固定的。

英文摘要

Clinical measurements are recorded for some patients but not others, at rates that differ across hospitals, and marginal conformal coverage does not ensure coverage within groups defined by missingness. We propose a missingness-aware conformal calibration procedure for mortality prediction under cross-hospital distribution shift. It selects a measurement on an independent sample, groups patients by whether that measurement is recorded, and applies Mondrian calibration within each group, so no calibration outcome is reused. We evaluate the procedure across hospitals in eICU and across care units within one MIMIC-IV hospital, using three predictors. Relative to pooled calibration, it reduces the average worst-group coverage gap on its selected groups in all six settings, with a median reduction of 1.9 percentage points; paired site-bootstrap intervals exclude zero in five. These gains do not extend uniformly. Calibration by predicted risk achieves smaller gaps on a broader panel of missingness groups, and when eICU hospitals are evaluated separately, the gain shrinks for all three predictors and reverses in sign for one. We explain this discrepancy with a hospital-level decomposition. Pooling reweights hospitals through a covariance between group shares and coverage errors, and lets errors of opposite sign cancel: weighting explains the reversal, and cancellation accounts for most of the attenuation for the other two predictors. Constructed population distributions show that pooled and within-hospital evaluations can rank calibration methods oppositely even without sampling noise. Pooled improvement alone therefore cannot establish better coverage within hospitals, even when the calibration groups are fixed.

Comments33 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑