用于CT理解的基于器官条件模式令牌的细粒度视觉语言预训练
Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding
AI总结:
研究从CT扫描和报告进行视觉语言预训练的难题,提出OCP-CT框架,通过保留全局对比分支、引入器官模式接口等方法,在公开基准上实现零样本异常诊断,相比先前结果有显著AUROC增益。
AI中文摘要:
从配对的CT扫描和放射学报告进行视觉语言预训练是一项可扩展但具有挑战性的任务。现有方法通常采用全局扫描-报告对比,可扩展但模糊了异质器官证据。同时,直接器官级对齐仍然粗糙。因此,预训练需要更精细的对齐单元:器官条件放射学模式。本文提出OCP-CT,一种用于CT视觉语言预训练的器官条件模式令牌对齐框架。具体包括保留稳定的全局CT-报告对比分支,引入器官模式接口,通过稀疏专家混合根据潜在放射学模式路由图像和文本令牌,可学习插槽将路由令牌查询为连续模式令牌,配对令牌对比将图像-文本模式令牌与基于报告衍生临床相似性构建的结构化软目标对齐。在公开可用的CT-RATE和RAD-ChestCT基准上,OCP-CT在零样本异常诊断中分别实现了84.5%和69.9%的平均AUROC。与先前最强结果相比,绝对AUROC增益分别为6.7和0.8个百分点。
英文摘要:
Computed tomography (CT) vision-language pretraining from paired volumes and radiology reports is a scalable yet challenging task. Existing methods commonly adopt global scan-report contrast, which is scalable but obscures heterogeneous organ evidence. Meanwhile, direct organ-level alignment remains coarse, since the same anatomy can exhibit multiple distinct radiological appearances. Therefore, pretraining requires a finer alignment unit: the organ-conditioned radiological pattern. In this work, we propose OCP-CT, an organ-conditioned pattern-token alignment framework for CT vision-language pretraining. Specifically, OCP-CT preserves a stable global CT-report contrastive branch and introduces an organ pattern interface: sparse Mixture-of-Experts (MoE) routes image and text tokens according to latent radiological patterns, learnable slots query the routed tokens into continuous pattern tokens, and paired token contrast aligns image-text pattern tokens with structured soft targets built from report-derived clinical similarity. On the publicly available CT-RATE and RAD-ChestCT benchmarks, OCP-CT achieves average AUROCs of 84.5% and 69.9% for zero-shot abnormality diagnosis, respectively. Compared with the strongest prior reported results, these results yield absolute AUROC gains of 6.7 and 0.8 percentage points.