arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解耦与推理:3D胸部CT中自由文本发现的解剖学引导两阶段体素级定位

Decouple and Reason: Anatomically Guided Two-Stage Voxel-Level Grounding of Free-Text Findings in 3D Chest CT

Kwang-Hyun Uhm, Inhwa Son, Sung-Jea Ko

arXiv 2607.12602首次发表:更新:

发表机构

Department of Artificial Intelligence, Gachon University; MEDAI(韩国加图立大学人工智能系; MEDAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究3D胸部CT中自由文本发现的体素级定位难题,提出解耦框架,分病变分割和文本-体积推理两阶段,利用解剖学引导解决空间模糊性,在基准测试中取得领先,证明解耦是处理该复杂性的有效范式。

AI 中文摘要

3D胸部计算机断层扫描(CT)中自由文本发现的自动体素级定位对临床可解释性至关重要。然而,由于大型3D体积的复杂空间复杂性和自由文本发现的异质性,这项任务仍然极具挑战性。现有的端到端方法往往难以同时学习准确3D分割所需的局部特征表示和文本对齐所需的复杂语义理解,导致定位性能次优。为克服这一基本限制,我们提出了一种新颖的解耦框架,将问题分解为两个专门阶段:(1)类别无关的病变分割和(2)文本-体积推理。这种结构分离使模型能够首先通过定位潜在异常来提取候选子体积。随后,进行密集的跨模态推理,将这些局部子体积与自由文本医学发现对齐。为解决局部区域固有的空间模糊性,推理模块通过利用相对空间坐标和肺叶先验进行显式解剖学引导。在ReXGroundingCT基准上评估,我们的方法在官方排行榜上的整体定位质量方面取得了领先水平。这些结果表明解耦检测与推理是处理3D医学视觉定位复杂性的高效范例。我们的代码可在此https URL公开获取。

英文摘要

Automatic voxel-level grounding of free-text findings in 3D chest Computed Tomography (CT) is critical for clinical interpretability. However, this task remains highly challenging due to the intricate spatial complexity of large 3D volumes and the heterogeneity of free-text findings. Existing end-to-end approaches often struggle to simultaneously learn the localized feature representations required for accurate 3D segmentation and the complex semantic understanding needed for text alignment, leading to suboptimal grounding performance. To overcome this fundamental limitation, we propose a novel decoupled framework that disentangles the problem into two specialized stages: (1) class-agnostic lesion segmentation and (2) text-volume reasoning. This structural separation allows the model to first extract candidate sub-volumes by localizing potential abnormalities. Subsequently, intensive cross-modal reasoning is performed to align these localized sub-volumes with free-text medical findings. To resolve the spatial ambiguities inherent in local regions, the reasoning module is augmented with explicit anatomical guidance, utilizing relative spatial coordinates and lung lobe priors. Evaluated on the ReXGroundingCT benchmark, our method achieves state-of-the-art performance in overall grounding quality on the official leaderboard. These results demonstrate that decoupling detection from reasoning is a highly effective paradigm for handling the complexity of 3D medical visual grounding. Our code is publicly available at https://github.com/khuhm/DAGG.

CommentsAccepted to MICCAI 2026 (Spotlight Presentation)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑