arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21197cs.LGq-bio.QMstat.ME

基于可靠性中心的稀疏纵向CT病灶大小预测评估:结合共形区间校准与Gompertz启发正则化

Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization

  • Vanderbilt University(范德堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Lingfei Kong

AI总结:

本研究基于稀疏纵向CT数据,构建基准并评估多种预测方法,发现额外历史观测收益有限,且准确性、可靠性与轨迹一致性需联合评估。

AI中文摘要:

稀疏的纵向CT随访限制了在仅有少量先前观测时对病灶大小的预测。我们从DeepLesion和Deep Lesion Tracker (DLT)构建了一个五次访视的DLT派生同病灶轨迹基准,共获得来自129名患者的205条轨迹。我们比较了一种探索性的常规稀疏到最终分析,与一种主要的固定访视索引水平线设计,该设计预测从T3到T4的常见对数变化,同时逐步添加更早的观测,评估预测准确性、不确定性可靠性、事后共形区间校准、亚组性能以及Gompertz启发的轨迹正则化。所评估的方法在点预测准确性上部分重叠,但不确定性行为不同。在十个训练种子上的平均留出集均方根误差(RMSE)在m = 1、2、3、4时分别为0.4726、0.4305、0.4499和0.4513,表明在m = 2时平均RMSE最低;额外的历史信息并未改善RMSE。在m = 4时,原始队列级特征高斯过程覆盖率接近95%的名义水平,而MC Dropout、深度集成和残差尺度区间则较为保守。患者级共形校准通常产生接近名义或保守的覆盖率,但代价是区间更宽。患者分组的开发交叉验证为Gompertz启发项选择了lambda* = 0。全局人群参考经常与病灶级变化方向相悖,预测难度在不同解剖亚组间存在差异。总体而言,一旦预测水平线被控制,额外的历史观测提供的预测收益有限,而预测准确性、不确定性可靠性和轨迹一致性并不一定同时改善,应在稀疏纵向成像中联合评估。

英文摘要:

Sparse longitudinal CT follow-up limits lesion-size forecasting when only a few prior observations are available. We constructed a five-visit DLT-derived same-lesion trajectory benchmark from DeepLesion and Deep Lesion Tracker (DLT), yielding 205 trajectories from 129 patients. We compared an exploratory conventional sparse-to-final analysis with a primary fixed visit-index horizon design predicting the common log change from T3 to T4 while progressively adding earlier observations, evaluating predictive accuracy, uncertainty reliability, post-hoc conformal interval calibration, subgroup performance, and Gompertz-inspired trajectory regularization. The evaluated methods showed partially overlapping point-prediction accuracy but distinct uncertainty behavior. Mean held-out RMSE across ten training seeds was 0.4726, 0.4305, 0.4499, and 0.4513 for m = 1, 2, 3, 4, indicating the lowest mean RMSE at m = 2; additional history did not improve RMSE. At m = 4, raw Cohort-Level Feature GP coverage was near the 95% nominal level, whereas MC Dropout, Deep Ensemble, and residual-scale intervals were conservative. Patient-level conformal calibration generally produced near-nominal or conservative coverage at the cost of wider intervals. Patient-grouped development cross-validation selected lambda* = 0 for the Gompertz-inspired term. A global population reference frequently opposed lesion-level change directions, and prediction difficulty varied across anatomical subgroups. Overall, additional historical observations provided limited predictive benefit once the prediction horizon was controlled, while predictive accuracy, uncertainty reliability, and trajectory consistency did not necessarily improve together, and should be evaluated jointly in sparse longitudinal imaging.

补充信息

↑