arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用器官分层报告知识学习基于解剖学的CT视觉语言表征

Learning Anatomy-Grounded CT Vision-Language Representations with Organ-Hierarchical Report Knowledge

Guoliang You, Hongming Li, Yuanwang Zhang, Yong Fan

arXiv 2607.10953首次发表:更新:

发表机构

Department of Radiology, Perelman School of Medicine, University of Pennsylvania(放射科,佩尔曼医学院,宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究利用CT图像与放射学报告进行医学视觉语言预训练,提出OKA-CT框架,通过转化报告为器官条件知识,分阶段学习,在数据集上实现零样本异常诊断高AUROC,优于基线,提升报告-图像对齐。

AI 中文摘要

医学视觉语言预训练(VLP)可实现可扩展的表征学习,但现有方法存在不足。我们提出OKA-CT,一个用于CT报告VLP的器官分层知识增强框架。它先将自由文本报告转化为器官条件知识,分两个学习阶段。阶段1通过细粒度器官条件监督注入解剖学证据,阶段2用器官特定报告证据指导对比学习。在CT-RATE和RAD-ChestCT数据集上,OKA-CT实现了零样本异常诊断AUROC分别为84.9和72.2,优于基线,还改善了报告-图像对齐。

英文摘要

Medical vision-language pretraining (VLP) from paired CT images and radiology reports enables scalable representation learning, but most existing methods align either whole scans with entire reports or local image regions with text fragments. These formulations underuse a key property of radiology reports: findings are organized around anatomical structures, with abnormalities described by organs, disease concepts, locations, and severity-related attributes. We propose OKA-CT, an organ-hierarchical knowledge-augmented framework for CT-report VLP. OKA-CT first converts free-text reports into organ-conditioned knowledge using radiology report parsing and LLM-assisted semantic structuring. The extracted hierarchy is used across two learning stages. Stage~1 injects anatomy-grounded evidence into the CT visual representation through fine-grained organ-conditioned supervision, while Stage~2 uses organ-specific report evidence to guide structured report-CT contrastive learning, where hierarchy-derived semantic soft targets treat non-paired cases with shared organ-level findings as weak semantic positives rather than uniform negatives. A lightweight query-based global branch further aggregates disease-relevant volumetric evidence for whole-scan representation. On CT-RATE and RAD-ChestCT datasets, OKA-CT achieves zero-shot abnormality diagnosis AUROCs of 84.9 and 72.2, outperforming prior CT VLP baselines. Retrieval and patch-occlusion analyses further show improved report-image alignment and stronger sensitivity to disease-associated anatomical regions.

Comments9 pages, 6 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑