arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大的、明亮的还是不可见的:3D CT基础模型的冻结特征基准测试

Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models

Maulik Chevli, Johannes Brandt, Rickmer Braren, Daniel Rueckert, Philip Müller

arXiv 2608.05960首次发表:更新:

发表机构

Technical University of Munich (TUM); TUM University Hospital; Imperial College London; Munich Center for Machine Learning (MCML); UKE Hamburg(慕尼黑工业大学(TUM); 慕尼黑工业大学医院; 伦敦帝国理工学院; 慕尼黑机器学习中心; 汉堡大学医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对10个冻结CT编码器开展基准测试,发现性能主要取决于异常的对比度与空间范围,小型低对比度病变仍是挑战,需区域或病变级预训练。

AI 中文摘要

常规CT解读本质上是综合性的,会捕获整个扫描范围内的偶然发现。3D CT基础模型可通过提供解剖结构和病理的通用表示来辅助这一过程。为评估其诊断广度,我们在三个胸部CT扫描队列(包括一个未见过的内部临床数据集)上,使用k近邻、零样本提示和线性探测对10个冻结CT编码器进行基准测试。我们发现不存在通用的最先进模型,排名会根据评估环境显著波动。虽然结合细粒度图像分词与视觉-语言对齐的模型通常表现最佳,但轻量级监督编码器仍具有很强竞争力,表明显式标签可有效替代规模。关键的是,我们观察到性能的主要决定因素不是模型架构,而是一个物理瓶颈:发现的可检测性与其与周围组织的对比度及空间范围成正比。通过受控的器官内比较,我们实证证明广泛存在或高对比度的异常(如装置和积液)可被可靠检测,而小型、低对比度的局灶性病变在所有评估的编码器中仍是持续存在的挑战。我们将此归因于全局池化嵌入的固有局限性,这表明要准确表示小型、低对比度结构,需要区域或病变级别的预训练。

英文摘要

Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models could assist this process by providing generalizable representations of anatomy and pathology. To evaluate their diagnostic breadth, we benchmark ten frozen CT encoders across three cohorts of thoracic CT scans, including an unseen internal clinical dataset, using $k$-nearest neighbors, zero-shot prompting, and linear probing. We find no universal state-of-the-art, with rankings fluctuating significantly depending on the evaluation context. While models combining fine-grained image tokenization with vision-language alignment generally perform best, a lightweight supervised encoder remains highly competitive, demonstrating that explicit labels can effectively substitute for scale. Crucially, rather than model architecture, we observe that the primary determinant of performance is a physical bottleneck: a finding's detectability scales with its contrast against surrounding tissue and its spatial extent. Through controlled within-organ comparisons, we empirically demonstrate that widespread or high-contrast abnormalities, such as devices and effusions, are reliably recovered. Conversely, small, low-contrast focal lesions remain a persistent challenge across all evaluated encoders. We attribute this to the inherent limitations of globally pooled embeddings, suggesting that accurately representing small, low-contrast structures will require region- or lesion-level pretraining.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑