发表机构
Center for Data Science, New York University; University of Chicago; NYU Grossman School of Medicine; Ataraxis AI(纽约大学数据科学中心; 芝加哥大学; 纽约大学格罗斯曼医学院; Ataraxis人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究生存模型中行政审查的陷阱,通过刻画行政截断泄漏发生的情况,将其与其他情况区分,提出检测方法,经模拟和真实队列验证,为生存预测提出数据集应提供足够潜在随访的设计原则。
AI 中文摘要
生存模型可利用部分观测数据对事件发生时间结果进行建模,广泛应用于临床预测。近期模型常依赖特定临床接触时收集的丰富输入数据,在大型回顾性数据集中,这些输入数据跨越多年,可能含获取时间线索。当仅在固定研究结束日期观察结果时会产生潜在失败模式,即行政截断泄漏。本文刻画了泄漏何时会发生,将其与经典信息审查和真实风险的时间变化区分开来,并提出检测方法。模拟表明行政截断泄漏会使固定水平AUC膨胀并影响Harrell's C指数,在真实乳房X光摄影队列中也有相同表现。这些结果为生存预测提出简单设计原则:对于n年预测任务,数据集应在最新输入日期后提供至少n年的潜在随访,否则模型可能受行政截断泄漏导致的偏差影响。
英文摘要
Survival models can model time-to-event outcomes using partially observed data. They are widely used in clinical prediction, including cancer risk, disease progression, treatment response, and mortality. Recent models often rely on rich inputs collected at a specific clinical encounter, such as medical images, laboratory tests, electronic health record snapshots, or sensor measurements. In large retrospective datasets, these inputs are usually collected over many calendar years. As a result, they may contain clues about when they were acquired through changes in devices, protocols, documentation, patient mix, or clinical practice. This creates a potential failure mode when outcomes are observed only up to a fixed study end date. More recent records necessarily have less potential follow-up than older records. A model that can infer the record date from the input may therefore learn to predict how much follow-up was available rather than the patient's true risk of experiencing the event. We call this failure mode administrative-cutoff leakage. In this paper, we characterize when this leakage can occur, distinguish it from classical informative censoring and genuine temporal changes in risk, and propose practical ways to detect it. In simulations, we show that administrative-cutoff leakage can inflate fixed-horizon AUC and can also affect Harrell's C-index under realistic follow-up patterns. We then demonstrate the same behavior in a real mammography cohort. These results motivate a simple design principle for survival prediction: for an n-year prediction task, the dataset should provide at least n years of potential follow-up after the latest input date. Otherwise, the models may be subject to bias induced by administrative-cutoff leakage.