发表机构
Université Laval; Radiation Oncology Service, Department of Specialized Medicine, Centre Intégré de Cancérologie (CIC), Hôpital de l’Enfant-Jésus, Centre Hospitalier Universitaire (CHU) de Québec-Université Laval; Service of Radio-oncology, CHU de Québec-Université Laval; Oncology division, Research Center of the CHU de Québec-Université Laval; Université Laval, Faculté de médecine(拉瓦尔大学; 魁北克大学附属医院(CHU)儿童耶稣医院综合癌症中心专科医学部放射肿瘤服务; 魁北克大学附属医院放疗服务; 魁北克大学附属医院研究中心肿瘤科; 拉瓦尔大学医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究验证了深度学习自动分割前列腺轮廓在LDR近距离放疗中剂量学上与专家手动勾画在群体层面等效,但个体等效需体积依赖阈值;DVH指数对IPSS预测无增量价值。
AI 中文摘要
背景与目的:剂量体积直方图(DVH)指数仍是毒性预测的主要剂量效应指标,但它们依赖于轮廓的勾画方式。观察者间变异性(IOV)不可避免且在临床上被接受,因此自动分割的问题在于等效性:自动轮廓产生的DVH误差是否与人类IOV相当,这些指数能否预测患者报告的毒性?材料与方法:在429例接受碘-125低剂量率近距离放疗(LDR-BT)单药治疗的患者中,将专家手动勾画的指数与确定性和贝叶斯nnUNet在固定剂量分布上的结果进行比较,采用双单侧检验,边界基于CT轮廓勾画的IOV设定。逻辑回归给出了90%等效概率下的患者层面阈值。在380例患者中,将11个DVH指数加入临床基线模型,用于预测国际前列腺症状评分(IPSS)在5年内六个时间点的变化,采用嵌套交叉验证、bootstrap区间和校正t检验。结果:所有队列层面的DVH指数比较均被宣布为等效,没有区间消耗超过其等效边界的44%。个体一致性较弱,等效率从55.9%到89.7%不等,取决于DVH指数和自动分割模型。阈值范围从0.864到0.958 Dice。在任何时间点,DVH模块均未改善IPSS预测。bootstrap区间宣布的最大改善为0.12个IPSS点,远低于最小临床重要差异,在所有学习器中均如此。结论:自动轮廓在平均意义上与专家剂量学在人类IOV范围内匹配,但个体等效性需要基于体积的质量指标阈值定义。我们的DVH指数面板对IPSS无显著预测能力,与分割来源和时间点无关。
英文摘要
Background and purpose: Dose Volume Histogram (DVH) indices remain the main dose-effect metrics for toxicity prediction, but they depend on how contours were made. Inter-observer variability (IOV) is unavoidable and clinically accepted, so the question for automatic segmentation addresses equivalence: do automatic contours produce DVH errors comparable to human IOV, and can these indices predict patient-reported toxicity? Materials and methods: In 429 patients treated with iodine-125 LDR-BT monotherapy, indices from expert manual delineation were compared with a deterministic and a Bayesian nnUNet on a fixed dose distribution, by two-one-sided tests against margins set from CT contouring IOV. Logistic regression gave patient-level thresholds at 90\,\% probability of equivalence. In 380 patients, eleven DHV indices were added to a clinical baseline predicting change in International Prostate Symptom Score (IPSS) at six horizons over 5 years, with nested cross-validation, bootstrap intervals and corrected $t$-tests. Results: All cohort-level comparisons of DVH indices were declared equivalent, with no interval consuming more than 44\,\% of its equivalence margin. Individual agreement was weaker with equivalence rates from 55.9\,\% to 89.7\,\%, depending on the DVH index and automatic segmentation model. Thresholds ranged from 0.864 to 0.958 Dice. No DVH block improved IPSS prediction at any horizon. The largest improvement declared by bootstrap intervals was 0.12 IPSS points, well below the minimal clinically important difference, across all learners. Conclusion: Automatic contours matched expert dosimetry within human IOV on average, but individual equivalence requires a volume dependent quality metric threshold definition. Our DVH indices panel provides no significant predictive power for IPSS, irrespective of segmentation source and horizon.