arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24486eess.IVcs.CV

模型效应还是标注效应?用于肺栓塞分割的优化标注及人类参照基准

Unified-protocol voxel-level pulmonary embolism annotations for three public CT angiography datasets

Qihang Sun, Zhongxiao Liu, Bailiang Jian, Shenman Qiu, Jingyuan Wang, Lei Zhang, Lixiang Xie, Jiazhen Pan, Christian Wachinger

首次发表
浏览论文内容

中文总结 AI 辅助

该研究量化了肺栓塞分割中评估标注与模型训练对性能的影响,发现标注效应至少与模型效应相当,建立了人类参照的nnPE基准模型及公开框架。

中文摘要 AI 辅助

目的:量化评估标注对肺栓塞(PE)分割性能测量的影响,对比模型训练变化的影响,并建立人类参照框架。材料与方法:本回顾性研究筛选了来自CADPE(91例)、FUMPE(35例)和READ(40例)的166个体素标注CT肺血管造影病例,最终纳入149例。一名主评估者按方案标注PE,一名高级胸放射科医生审核并修正所有分割结果;三家机构的三名额外评估者对15例子集进行标注。通过对比两个预训练nnU-Net模型(nnU-Net-A、nnU-Net-B)在原始标注与优化标注上的表现,测量标注效应;在标注固定的情况下,对比相同架构在不同数据集组合上的训练结果,测量模型效应。基准模型(nnPE)采用留一数据集法及合并五折交叉验证训练,分析四类指标,采用配对Wilcoxon符号秩检验、Benjamini-Hochberg校正及自助法95%置信区间。结果:仅改变标注使nnU-Net-A的平均DSC提升0.143(95%CI:0.122-0.166),nnU-Net-B提升0.188(95%CI:0.163-0.213),两者P均<0.001;而改变训练数据集组合使DSC变化0.028。在CADPE和FUMPE上标注效应超过模型效应,READ上为0.045。重新标注后,三个数据集的掩码内衰减标准差均下降,P均<0.001。nnPE在合并交叉验证中DSC达0.72±0.22,但在52次配对对比中得分低于所有四名评估者,校正后P均<0.05。结论:评估标注对PE分割性能测量的影响至少与模型训练选择相当,人类参照评估框架已公开供未来研究使用。

英文摘要

Reliable clot-volume quantification and subsequent risk assessment in pulmonary embolism depend on precise segmentation of emboli on computed tomography pulmonary angiography. Deep learning models for this task must be trained on accurate voxel-level labels. The three public datasets that provide such labels were annotated under different protocols, and some of their studies contain unlabeled emboli or labels that are discontinuous across slices. This Data Descriptor presents voxel-level pulmonary embolism annotations for 149 of the 166 studies in these datasets. A primary rater drew all annotations under a single protocol. A thoracic radiologist with more than 20 years of experience reviewed and revised them. Three raters at three different centers independently annotated a subset of 15 studies. The subset was selected by source dataset and embolus location. Technical validation quantifies volumetric agreement with the source annotations, changes in within-mask attenuation, and inter-rater agreement on the subset. The dataset is intended to allow segmentation models to be developed and compared under a common reference standard.

发表机构

  • Technical University of Munich (TUM)(慕尼黑工业大学)
  • Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
  • The Affiliated Hospital of Xuzhou Medical University(徐州医科大学附属医院)
  • The Third Affiliated Hospital of Soochow University(苏州大学附属第三医院)
  • The Affiliated Taizhou People’s Hospital of Nanjing Medical University(南京医科大学附属泰州人民医院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑