arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28455cs.CVcs.AI

ARC-CT:用于3D胸部CT的解剖路由对比视觉-语言学习

ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT

Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol, Şeyda Ertekin

首次发表
浏览论文内容

中文总结 AI 辅助

ARC-CT是一种基于LLM提取的报告标签、无需人工标注的区域感知对比视觉-语言框架,通过三个组件解决胸部CT全局对比学习的局限,在18种异常上实现0.86无掩码宏观AUC,性能优于高效基线和大Transformer模型。

中文摘要 AI 辅助

对比视觉-语言学习利用配对的胸部CT体积和放射学报告,无需人工标注标签即可学习异常分类器。然而,胸部CT的两个特性对传统的全局对比学习构成挑战:其一,许多关键异常较小或在解剖学上具有局限性,将整个体积池化为单个嵌入可能会稀释其视觉证据;其二,标准对比目标将批次中的每一次其他扫描视为负样本,由于许多胸部CT存在共同异常,该目标会错误地将同正样本对推开。我们提出用于3D胸部CT的解剖路由对比学习(ARC-CT),这是一种仅使用大型语言模型从报告中提取的标签、无需人工标注或边界框即可解决这些限制的区域感知框架。ARC-CT包含三个组件:(1)AnatomyQFormer,通过自动生成的器官掩码约束的查询来定位证据;(2)标签-Jaccard软InfoNCE目标,将标准独热目标与每对的标签集重叠相结合,减少存在共同临床发现的研究之间的假阴性惩罚;(3)器官级对齐损失,将掩码池化的视觉特征与离线用大型语言模型提取的器官特异性报告文本相连接。ARC-CT使用紧凑的3D ResNet-18主干,在18种异常上实现了0.86的无掩码宏观AUC,总体而言,ARC-CT优于可比的高效基线和几个更大的Transformer模型,其代码和权重可在指定的URL获取。

英文摘要

Contrastive vision-language learning uses paired chest CT volumes and radiology reports to learn abnormality classifiers without manually annotated labels. However, two characteristics of chest CT challenge conventional global contrastive learning. First, many critical abnormalities are small or anatomically localized, and pooling an en- tire volume into a single embedding may dilute their visual evidence. Second, the standard contrastive objective treats every other scan in a batch as a negative. Because many chest CTs share abnormalities, this objective incorrectly pushes co-positive pairs apart. We propose Anatomy-Routed Contrastive Learning for 3D Chest CT (ARC-CT), a region-aware framework that addresses these limitations using only la- bels extracted from reports by an LLM, with no manual annotations or bounding boxes. ARC-CT combines three components: (1) an Anato- myQFormer localizing evidence via queries constrained by automatically generated organ masks; (2) a label-Jaccard soft InfoNCE objective in- tegrating the standard one-hot target with the label-set overlap of each pair, which reduces false-negative penalties between studies that share clinical findings; and (3) an organ-level alignment loss connecting mask- pooled visual features to organ-specific report text extracted offline with a large language model. ARC-CT achieves a 0.86 mask-free macro AUC across 18 abnormalities using a compact 3D ResNet-18 backbone. Over- all, ARC-CT outperforms both comparable efficient baselines and sev- eral larger transformer models. Our code and weights are available at https://github.com/arc-ct/arc-ct.

发表机构

  • Harvard Medical School(哈佛医学院)
  • METU-DTX Digital Transformation and Innovation Center(METU-DTX数字转型与创新中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑