arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

天花板在通道中:临床预测中的学习者差距与测量前沿审计

The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction

Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan

arXiv 2609.01909首次发表:更新:

发表机构

Luddy School of Informatics, Computing, and Engineering, Indiana University Indianapolis; Ahsanullah University of Science and Technology; Purdue University Indianapolis; Indiana University School of Medicine(印第安纳大学印第安纳波利斯分校勒迪信息学、计算与工程学院; 阿萨努拉科技大学; 普渡大学印第安纳波利斯分校; 印第安纳大学医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出区分临床预测饱和原因的学习者差距与测量通道天花板概念,验证其在三大真实队列中有效,为临床预测饱和提供可审计决策框架。

AI 中文摘要

临床预测的饱和现象源于两种不同原因:拟合的学习器可能无法提取可用信息,或记录的变量会施加群体前沿。我们通过学习者差距(learner gap)和测量通道天花板(measurement-channel ceiling)区分这两个量。最优平衡准确率由总变差分离表征,由此得出架构不变性、替换污染下的尖锐部分识别结果、交叉拟合天花板估计量以及多模态决策改进的精确条件。我们添加两个有限样本诊断指标:标签置换乐观度下限和欠拟合曲线,并在三个真实队列中验证该审计方法:UCI再入院队列(样本量n=99343)、BRFSS糖尿病队列(样本量n=253680)和NHANES糖化血红蛋白(HbA1c)队列(样本量n=10219)。调优良好的梯度提升模型在UCI和BRFSS队列中几乎达到估计的前沿,而刻意或实际存在缺陷的学习器则保留较大差距。NHANES队列显示问卷与测量的边际前沿无差异,但存在显著的联合互补增益,修正了“客观模态必须占优”的简单论断。在所有队列中,适度的AUROC增益与大得多的贝叶斯决策翻转率共存,多个架构估计出相似的前沿,但其达到的平衡准确率却差异显著。随后,对104项临床任务的PRISMA引导综合分析显示,相同的通道级规律在超过18种疾病类别中重复出现:广泛但非通用的结构化临床区域、模型家族内相同通道的增益递减、测量通道改变时性能更高。该框架将饱和现象从经验观察转化为可审计的决策:存在提升空间时改进学习器;否则改进测量。

英文摘要

Clinical prediction can saturate for two different reasons: a fitted learner may fail to extract available information, or the recorded variables may impose a population frontier. We separate these quantities through the \emph{learner gap} and the \emph{measurement-channel ceiling}. Optimal balanced accuracy is characterized by total-variation separation, yielding architecture invariance, a sharp partial-identification result under replacement contamination, a cross-fitted ceiling estimator, and exact conditions for multimodal decision improvement. We add two finite-sample diagnostics, namely a label-permutation optimism floor and an underfit curve, and validate the audit on three real cohorts: UCI readmission ($n=99{,}343$), BRFSS diabetes ($n=253{,}680$), and NHANES HbA1c ($n=10{,}219$). Well-tuned gradient boosting nearly reaches the estimated frontier in UCI and BRFSS, whereas deliberately or practically deficient learners retain large gaps. NHANES yields a null difference between questionnaire and measured marginal frontiers but a significant joint complementarity gain, refining the simplistic claim that an objective modality must dominate. Across all cohorts, modest AUROC gains coexist with substantially larger Bayes decision-flip rates, and several architectures estimate similar frontiers while their achieved balanced accuracy differs sharply. A PRISMA-guided synthesis of 104 clinical tasks then shows that the same channel-level regularities recur across more than 18 disease categories: a broad but non-universal structured-clinical region, diminishing same-channel gains across model families, and higher performance when measurement channels change. The framework converts saturation from an empirical observation into an auditable decision: improve the learner when headroom remains; improve measurement when it does not.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑