arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

腹部CT中三维视觉定位的规模化

Scaling 3D Visual Grounding in Abdominal CT

Sam Church, Danyal Maqbool, Joshua D. Warner, Andrew Voter, Junjie Hu, Meghan G. Lubner, Tyler J. Bradshaw

arXiv 2610.04095首次发表:更新:

发表机构

University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出利用常规PACS二维标注自动生成大规模三维视觉定位训练数据,并构建LocusCT模型,在腹部CT基准上显著优于对比方法。

AI 中文摘要

视觉定位模型通过将报告中的发现与图像区域关联起来,可以增强放射学工作流程。这对于三维CT尤其有价值,因为发现往往只占据体积的一小部分。训练三维定位模型需要大量配对的短语和区域,而构建此类数据集成本高昂,需要放射科医生手动标注图像。我们假设这种监督在常规报告过程中已经隐式产生,因为放射科医生经常在关键图像上放置二维标注(例如距离测量、箭头)以进行测量并支持报告解读。我们引入了一种自动化流程,将这些常规临床标注转化为用于三维视觉定位的大规模短语-区域监督。该流程通过元数据匹配将每个标注与报告中的相应发现关联起来,然后使用可提示的三维分割模型将二维标注转换为体积掩膜。这无需额外的放射科医生标注即可生成短语-掩膜-体积数据集。应用于单一机构的临床影像归档与通信系统(PACS)时,我们的方法从59,000次腹部CT检查中生成了105,000个短语-掩膜-体积三元组。我们还引入了两个腹部CT定位基准,LocusBench-Onc和LocusBench-ED,分别包含240个肿瘤学和260个急诊科放射科医生审核的短语-掩膜-体积三元组,后者涵盖阑尾炎、血肿和疝气等13个不同类别。我们进一步引入了LocusCT,一个在此数据集上训练的三维视觉定位模型,其在LocusBench-Onc上达到0.725的命中率,在LocusBench-ED上达到0.773,显著优于对比模型。这些结果表明,常规PACS标注是三维视觉定位的可扩展且此前未使用的监督来源。

英文摘要

Visual grounding models can enhance radiology workflows by linking report findings to image regions. This is particularly valuable for 3D CT, where findings often occupy a tiny fraction of the volume. Training 3D grounding models requires large sets of paired phrases and regions, and building such datasets is expensive, requiring radiologists to annotate images by hand. We posit that this supervision is already created implicitly during routine reporting, as radiologists frequently place 2D annotations (e.g., distance measurement, arrows) on key images to make measurements and to support report interpretation. We introduce an automated pipeline that converts these routine clinical annotations into large-scale phrase-region supervision for 3D visual grounding. The pipeline links each annotation to the corresponding finding in the report through metadata matching, then uses a promptable 3D segmentation model to convert the 2D annotation into a volumetric mask. This produces phrase-mask-volume datasets without requiring additional radiologist annotation. Applied to a single institution's clinical picture archiving and communication system (PACS), our approach generated 105K phrase-mask-volume triplets from 59K abdominal CT exams. We also introduce two abdominal CT grounding benchmarks, LocusBench-Onc and LocusBench-ED, which comprise 240 oncology and 260 emergency-department radiologist-reviewed phrase-mask-volume triplets, respectively, with the latter spanning 13 distinct categories such as appendicitis, hematoma, and hernia. We further introduce LocusCT, a 3D visual grounding model trained on this dataset, which achieves hit rates of 0.725 on LocusBench-Onc and 0.773 on LocusBench-ED, substantially outperforming comparator models. These results show that routine PACS annotations are a scalable, previously unused source of supervision for 3D visual grounding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑