arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

胸部CT中无文本的3D异常分割的实例引导报告锚定

Instance-Guided Report Anchoring for Text-Free 3D Abnormality Segmentation in Chest CT

Zhenyu Bu, Haoyan Ding, Chushu Shen, Xinyuan Zheng, Peiyu Duan, Xueqi Guo, Sepehr Farhand, Yoshihisa Shinagawa, Gerardo Hermosillo, Chaowei Wu

arXiv 2609.00447首次发表:更新:

发表机构

Siemens Medical Solutions USA, Inc.; The Ohio State University; University of California, Los Angeles; Biomedical Imaging Research Institute, Cedars-Sinai Medical Center; Yale University(西门子美国医疗解决方案公司; 俄亥俄州立大学; 加州大学洛杉矶分校; 西达赛奈医疗中心生物医学成像研究所; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出与模型无关的IGRA模块,通过训练时锚定异常实例与对应报告发现,推理时丢弃文本组件,将ReXGroundingCT的自由文本定位转为多标签体积分割,提升胸部CT无文本3D异常分割性能。

AI 中文摘要

胸部CT中准确的3D异常分割需要密集的空间监督,但获取专家体素级标签成本高昂。然而,放射学报告在临床解读过程中会常规生成,其中包含针对具体实例的描述,这些描述无需新的密集标注即可提供额外指导。现有视觉-语言定位方法通常需要在推理时使用报告生成的发现,这使得定位依赖于配对文本,且每次前向传播仅能处理一个查询到的发现。我们提出实例引导报告锚定(Instance-Guided Report Anchoring,IGRA),这是一种与模型无关的模块,可保留每个标注的异常实例与描述它的报告发现之间的对应关系。IGRA在训练期间池化每个实例表示并将其锚定到对应的发现嵌入;推理时会丢弃所有与文本相关的组件。我们还通过合并相同类别的实例,将ReXGroundingCT上的自由文本定位重新表述为多标签体积分割,从而允许在仅使用图像的单次前向传播中预测所有异常类别。IGRA相比最强的仅图像基线将Dice提升了22.5%(30.93 vs. 25.25),在单发现子集上与VoxTell相当(30.29 vs. 30.43)。将IGRA原封不动地应用于四个标准3D分割骨干网络,可在所有架构上提升Dice和命中率。在LIDC-IDRI、PleThora及一个内部私有数据集上的零样本评估进一步显示,相比仅图像基线获得了持续的提升。

英文摘要

Accurate 3D abnormality segmentation in chest CT requires dense spatial supervision, but obtaining expert voxel-level labels is costly. Radiology reports, however, are routinely generated during clinical interpretation and contain instance-specific descriptions that can provide additional guidance without new dense annotation. Existing vision-language grounding methods typically require report-derived findings at inference, making localization dependent on paired text and limiting each forward pass to a queried finding. We propose Instance-Guided Report Anchoring (IGRA), a model-agnostic module that preserves the correspondence between each annotated abnormality instance and the report finding that describes it. IGRA pools each instance representation and anchors it to the corresponding finding embedding during training; all text-related components are discarded at inference. We further reformulate free-text grounding on ReXGroundingCT as multi-label volumetric segmentation by merging same-category instances, allowing all abnormality categories to be predicted in one image-only forward pass. IGRA improves Dice by 22.5% over the strongest image-only baseline (30.93 vs. 25.25) and is comparable to VoxTell on the single-finding subset (30.29 vs. 30.43). Applied unchanged to four standard 3D segmentation backbones, IGRA improves Dice and hit rate across all architectures. Zero-shot evaluation on LIDC-IDRI, PleThora, and a private in-house dataset further shows consistent gains over image-only baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑