arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19526cs.LG

C 指数错觉:已发表生存模型中未经校准的歧视

The C-index illusion: discrimination without calibration in published survival models

  • Eastern University(东方大学)

机构由 AI 辅助整理,请以论文原文为准。

Rafael da Silva, Danilo Alvares

中文总结 AI 辅助

研究指出仅用 C 指数评估生存模型有误导性,通过在三个领域重现三个已发表模型验证。发现模型存在校准失败等问题,如概率估计随预测期下降等,还发布了评估工具包,揭示了模型评估中对 C 指数的错误信心问题。

中文摘要 AI 辅助

《评估生存分析模型时停止追逐 C 指数》(ICML 2026,焦点论文)基于合成数据从规范角度指出,仅通过歧视性指标(即一致性指数)评估生存模型会产生系统性误导的模型比较,因为该指标忽略了校准和时间依赖性准确性。对于真实已发表的非临床模型是否如此尚未得到验证。我们在三个结构不同的领域(硬盘故障、点对点信贷违约和数字平台上的用户脱离)重现了三个已发表的生存 - 机器学习模型,根据锚定论文自己的合成实验验证了我们的评估工具,并在霍尔姆校正的家族式错误率下测试了五个预先注册的假设。五个假设中有三个被拒绝。一个几乎完全重现已发表歧视性结果(C = 0.9595 对比报告的 0.958)的模型在 p = 2.6e - 136 时未能通过正式校准测试;广泛的特征消融搜索未发现导致这种歧视的单一属性,因此校准失败不是简单捷径造成的假象。当将贷款提前还款视为非信息性删失而非竞争风险时,贷款人估计的违约风险向上偏差约两个百分点,在风险最高的部分增长到近四个百分点。一个平台的流失模型显示,即使其全局歧视性保持在预先注册的 C 指数范围内,概率估计也会随着预测期而下降。对指标选择是否会颠倒首选模型的直接测试未被拒绝,不过由于每个领域只有两到三个模型,检验效力有限;我们记录的失败模式更应被描述为对所选模型的错误信心,而非选择了错误模型。我们发布了一个可重复使用的、预先注册的评估工具包,带有完整代码和一个带注释的笔记本。

英文摘要

Recent work has argued normatively, on synthetic data, that evaluating survival models by discrimination alone (concordance index) yields systematically misleading model comparisons, because the metric ignores calibration and time-dependent accuracy. Whether this matters for real, published, non-clinical models has not been tested. We reproduce three published survival-ML models across three structurally distinct domains -- hard-drive failure prediction, peer-to-peer credit default, and user disengagement on digital platforms -- validate our instrument against the anchor paper's own synthetic experiment, and test five pre-registered hypotheses under a Holm-corrected family-wise error rate. Three of five reject (though one pre-registered threshold clears by a narrow margin). A model reproducing the published literature's discrimination almost exactly (C = 0.9595 vs. 0.958 reported) fails a formal calibration test at p < 0.001; a broad feature-ablation search finds no single attribute responsible for its discrimination, so the calibration failure is not a trivial shortcut artifact. A lender's estimated default risk is biased upward by roughly two percentage points, growing to nearly four in the riskiest segment, when loan prepayment is treated as non-informative censoring rather than a competing risk. A platform's churn model shows probability estimates that degrade with the horizon even as global discrimination stays within the pre-registered C-index band. A direct test of whether metric choice inverts model preference does not reject, though with limited power given two to three models per domain; the failure mode we document is better characterized as misplaced confidence in a chosen model than as choosing the wrong one. We release a pre-registered evaluation harness with full code and an annotated notebook, so these results can be verified independently and the audit extended.

补充信息

↑