发表机构
Eindhoven University of Technology; Catharina Hospital Eindhoven(埃因霍温理工大学; 埃因霍温卡塔琳娜医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以对比增强CT影像的rPCI分割为对象,量化观察者间与nnU-Net模型的变异性,发现rPCI评分对典型分割变异性具鲁棒性,仅临界病例需专家审查。
AI 中文摘要
深度学习分割模型常采用Dice、HD95、ASD等几何指标进行评估,但这些指标的提升在多大程度上转化为下游决策的临床有意义变化仍不明确。本研究针对对比增强CT影像的放射学腹膜癌指数(rPCI)区域分割,探究该指标与决策间的差距:共识定义提供了解剖学基础的3D区域,临床使用的PCI 20阈值支持决策级评估。针对10例腹部CT扫描,量化4名专家的观察者间变异性;在全部13个区域上,采用Dice、HD95、ASD将已发表的基于nnU-Net的rPCI分割模型与人类参考基准进行对比。为关联几何差异与临床影响,在多数投票rPCI图谱上实施概率性腹膜转移模拟,将区域边界变异性传播至衍生(r)PCI评分及PCI 20临界值下的分类变异性。结果显示,观察者间一致性较高(平均Dice为0.87);模型在多数区域匹配人类表现,但在4、8区及小肠区域(9-12区)偏差更大。模拟中,观察者与模型的评分差异通常较小(平均ΔrPCI约0.3-0.6),决策翻转主要发生在参考评分接近20时。这些结果表明,基于rPCI的评分对典型分割变异性总体具有鲁棒性,同时强调临界病例是专家审查仍必不可少的主要场景。
英文摘要
Deep learning segmentation models are often evaluated using geometric metrics such as Dice, HD95, and ASD, yet it remains unclear to what extent improvements in these metrics translate into clinically meaningful changes in downstream decision-making. The metric-to-decision gap is examined using radiological Peritoneal Cancer Index (rPCI) region segmentation on contrast-enhanced CT, where a consensus definition provides anatomically grounded 3D regions and the clinically used PCI 20 threshold enables decision-level evaluation. Inter-observer variability is quantified across four experts on ten abdominal CT scans, and a published nnU-Net based rPCI segmentation model is benchmarked against this human reference using Dice, HD95, and ASD across all 13 regions. To relate geometric differences to clinical impact, a probabilistic peritoneal metastasis simulation is implemented on majority-vote rPCI maps, propagating region-boundary variability into variability of derived (r)PCI scores and classification at the PCI 20 cutoff. Observers showed high agreement (mean Dice $0.87$), while the model matched human performance in most regions but deviated more in regions 4, 8, and the small-bowel regions (9-12). Across simulations, score differences were typically small (mean $Δ$rPCI $\approx 0.3$-$0.6$) for both observers and the model, and decision flips occurred predominantly when the reference score was near 20. These results suggest that rPCI-derived scoring is generally robust to typical segmentation variability, while highlighting borderline cases as the main setting where expert review remains essential.
CommentsAccepted as a CaPTion workshop paper at MICCAI 2026