目标泄漏而非模型类别解释了基于调查的心血管筛查中报告的准确性:对玻璃盒模型和表格基础模型的泄漏分层审计
Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models
浏览论文内容
中文总结 AI 辅助
本研究通过泄漏分层审计发现,心血管筛查模型报告的准确性主要源于目标泄漏而非模型类别,玻璃盒模型在保持性能的同时支持公平性修复和高效推理。
中文摘要 AI 辅助
基于全国健康调查训练的心血管筛查模型通常报告受试者工作特征曲线下面积(AUROC)接近0.89。我们探究了这种准确性反映的是学习还是目标泄漏,表格基础模型是否改变了答案,以及部署所需的属性是否能经受联合检验。我们在2022年行为风险因素监测系统的442,067名受访者中,对跨越线性、树集成、神经网络、玻璃盒和表格基础类别的十种分类器进行了流行性心肌梗死基准测试,覆盖五个泄漏风险递减的特征层级。每个模型都接受了判别力、校准、在明确筛查阈值下的公平性、共形覆盖、解释忠实性和推理成本的审计,然后——模型和阈值冻结——应用于2023年的430,755名受访者。移除两个诊断后特征使每个模型损失0.049-0.051的AUROC,将整个领域压缩到0.0045宽的带内。玻璃盒可解释提升机在预先指定的0.005边际内不劣于所有替代方案,同时对队列评分速度约为最强基础模型的104倍。一个阈值检测到女性梗死的75.4%,而男性为89.0%;编辑模型的形状函数将差距缩小到0.010。边际共形预测对男性给出0.86的覆盖率,对60岁以上成年人给出0.82;蒙德里安校准修复了每个层级。冻结模型在0.002 AUROC内迁移。该文献中报告的性能余量是特征集的属性,而非学习器的属性。透明性没有带来可测量的成本,并使公平性修复和不确定性条件化直接可审计。评估实践,而非模型容量,是约束条件。
英文摘要
Cardiovascular screening models trained on national health surveys routinely report areas under the receiver operating characteristic curve (AUROC) near 0.89. We asked whether that accuracy reflects learning or target leakage, whether tabular foundation models change the answer, and whether the properties deployment requires survive joint examination. We benchmarked ten classifiers spanning linear, tree-ensemble, neural, glass-box, and tabular foundation classes for prevalent myocardial infarction in 442,067 respondents of the 2022 Behavioral Risk Factor Surveillance System across five feature tiers of decreasing leakage risk. Each was audited for discrimination, calibration, fairness at an explicit screening threshold, conformal coverage, explanation faithfulness, and inference cost, then applied -- models and thresholds frozen -- to 430,755 respondents of 2023. Removing two post-diagnostic features cost every model 0.049-0.051 AUROC, collapsing the field into a 0.0045-wide band. The glass-box explainable boosting machine was non-inferior to every alternative within a pre-specified 0.005 margin while scoring the cohort roughly 104 times faster than the strongest foundation model. One threshold detected 75.4% of women's infarctions against 89.0% of men's; editing the model's shape functions reduced the gap to 0.010. Marginal conformal prediction gave 0.86 coverage to men and 0.82 to adults over 60; Mondrian calibration repaired every stratum. Frozen models transported within 0.002 AUROC. Reported headroom in this literature is a property of the feature set, not the learner. Transparency cost nothing measurable and made fairness repair and uncertainty conditioning directly auditable. Evaluation practice, not model capacity, is the binding constraint.
发表机构
- XU Exponential University of Applied Sciences(XU指数应用科学大学)
- German University of Digital Science(德国数字科学大学)
- Abu Dhabi University(阿布扎比大学)
- Zayed University(扎耶德大学)
- Cleveland Clinic Hospital(克利夫兰诊所医院)
机构由 AI 辅助整理,请以论文原文为准。