arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CT基础模型的解剖学上下文适配

Anatomy Contextualized Adaptation of CT Foundation Models

Roshan Kenia, Stephanie L McNamara, William Lotter

arXiv 2607.27154首次发表:更新:

发表机构

Harvard Medical School; Dana-Farber Cancer Institute; Massachusetts General Hospital; Brigham and Women’s Hospital(哈佛医学院; 达纳-法伯癌症研究所; 麻省总医院; 布里格姆妇女医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出轻量框架ACA,适配冻结CT基础模型实现解剖学级视觉-语言对齐,在Merlin、CT-RATE数据集上零样本发现分类性能优于基线,训练耗时短且可保留增强全局解剖学上下文。

AI 中文摘要

CT视觉-语言基础模型在下游任务中展现出良好性能,但通常采用全体积表示,会稀释细粒度的解剖学信号。细粒度视觉-语言预训练通过将解剖学级视觉特征与特定解剖学术语文本对齐来解决该问题,但会丢失全体积模型提供的全局上下文,且现有细粒度方法从头开始训练,计算成本较高。我们提出解剖学上下文适配(Anatomy Contextualized Adaptation, ACA),这是一种轻量框架,用于适配冻结的CT基础模型表示,实现解剖学级视觉-语言对齐,同时增强全局上下文。ACA使用TotalSegmentator将CT体积分解为解剖学级嵌入,通过捕捉跨解剖学关系的Transformer对嵌入进行优化,并将其与从放射学报告中提取的每个解剖学及扫描级文本对齐。在Merlin和CT-RATE数据集上的评估显示,ACA在零样本发现分类任务中始终优于冻结基础模型基线和现有细粒度方法,且嵌入缓存后训练耗时不足1小时。ACA的跨解剖学Transformer学习到的注意力权重还显示出合理的跨解剖学上下文路由。总体而言,这些结果表明ACA是一种轻量方法,可在适配CT基础模型到基于解剖学的视觉-语言对齐的同时,保留并增强全局解剖学上下文。

英文摘要

CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight framework that adapts frozen CT foundation model representations for anatomy-level vision-language alignment while enhancing global contextualization. ACA uses TotalSegmentator to decompose CT volumes into anatomy-level embeddings, which are refined via a transformer that captures cross-anatomy relationships, and aligned to both per-anatomy and scan-level text extracted from radiology reports. Evaluated on Merlin and CT-RATE, ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification, while requiring less than one hour of training once embeddings are cached. The attention weights learned by ACA's inter-anatomy transformer additionally indicate plausible cross-anatomy context routing. Altogether, these results support ACA as a lightweight approach for adapting CT foundation models to anatomically grounded vision-language alignment while preserving and enhancing global anatomical context.

CommentsProceedings of the European Conference on Computer Vision (ECCV) 2026 MedFM-Bench Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑