arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04025stat.MEcs.LG

面向病理生存模型的感知删失共形下界预测边界的多队列验证

A Multi-Cohort Validation of Censoring-Aware Conformal Lower Predictive Bounds for Pathology Survival Models

Mingi Hong

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过多队列验证,评估了用于病理生存模型的感知删失共形下界预测边界方法drcosarc,发现其性能依赖于队列,解释受患者单元、估计量和删失假设影响。

中文摘要 AI 辅助

全切片生存模型通常仅提供风险排名,而无法对个体事件时间给出校准性表述。我们对固定截断的drcosarc进行评估,这是一种用于离散时间多实例学习生存头的事后共形包装器,采用冻结的UNI2-h表示,在内部的18种配置扫描中覆盖5个TCGA队列,并在外部的5种配置评估中覆盖3个CPTAC队列。我们区分了逆概率删失加权(IPCW)估计的配置-折-拆分汇总,以及中位数下界预测边界(LPB)与层次感知患者集合估计量的平均drcosarc-朴素LPB差值。在α=0.1时,drcosarc的IPCW估计在KIRC、LUAD和STAD中最接近0.90。患者集合的drcosarc-朴素区间在KIRC、KIRP、STAD、UCEC和CPTAC-CCRCC中排除零,但在内部LUAD、CPTAC-LUAD、CPTAC-UCEC以及内部LUSC扩展中包含零。在具有已知事件时间的20次重复低删失半合成设置中,drcosarc的经验覆盖率为0.9129 [0.9053, 0.9207]。一项探索性分析支持该数据生成过程中存在删失导致的头误差交互作用。在两队列ABMIL敏感性分析中,将风险网格增加到K=16会使局部边际IPCW估计超过预先设定的0.87阈值,并产生正的配对LPB差值,尽管最差组估计仍低于0.87。总体而言,性能依赖于队列,且其解释随患者水平单元、估计量和删失假设而变化。

英文摘要

Whole-slide survival models commonly provide risk rankings without calibrated statements about individual event times. We evaluate fixed-cutoff drcosarc, a post-hoc conformal wrapper for discrete-time multiple-instance learning survival heads using frozen UNI2-h representations, in an internal 18-configuration sweep across five TCGA cohorts and an external five-configuration evaluation across three CPTAC cohorts. We distinguish configuration--fold--split summaries of the inverse-probability-of-censoring-weighted (IPCW) estimate and median lower predictive bound (LPB) from a hierarchy-aware patient-ensemble estimand of the mean drcosarc--naive LPB difference. At $α=0.1$, the drcosarc IPCW estimate was nearest 0.90 in KIRC, LUAD, and STAD. Patient-ensemble drcosarc--naive intervals excluded zero in KIRC, KIRP, STAD, UCEC, and CPTAC-CCRCC, but included zero in internal LUAD, CPTAC-LUAD, CPTAC-UCEC, and the internal LUSC extension. In a 20-replicate low-censoring semi-synthetic setting with known event times, drcosarc empirical coverage was 0.9129 [0.9053, 0.9207]. An exploratory analysis supported a head-error-by-censoring interaction within that data-generating process. In a two-cohort ABMIL sensitivity analysis, increasing the hazard grid to $K=16$ raised localized marginal IPCW estimates above the prespecified 0.87 threshold and yielded positive paired LPB differences, although worst-group estimates remained below 0.87. Overall, performance was cohort dependent, and its interpretation changed with the patient-level unit, estimand, and censoring assumptions.

补充信息

↑