3D计算机断层扫描中视觉语言建模的统一监督
Unified Supervision For Vision-Language Modeling in 3D Computed Tomography
中文总结 AI 辅助
针对通用视觉语言模型在3D CT放射诊断中精度不足、标注数据异质稀缺的问题,提出统一多类监督信号的体积VLM Uniferum,实现SOTA性能与强泛化能力,为医学影像VLM发展指明新方向。
中文摘要 AI 辅助
通用视觉语言模型(VLM)已成为放射学领域颇具前景的工具,其具备的零样本能力可降低对大规模标注数据集的需求。然而在诊断放射学这类高风险领域中,这类模型往往缺乏可靠临床应用所需的判别精度。公开的体积CT数据集数量稀少且异质性强,标注格式和粒度差异极大,进一步加剧了这一挑战。为解决上述局限,我们提出Uniferum——一种体积VLM,可将分类标签和分割掩码中编码的各类监督信号统一到单一训练框架中。通过协调三个带有不同标注的公开3D CT数据集,Uniferum取得了当前最优性能,在CT-RATE基准上的AUROC相较于基于CLIP的模型和传统多标签卷积模型提升了7%。该模型展现出强大的分布外泛化能力,在RAD-CHEST和INSPECT数据集上观测到了出人意料的零样本性能表现。我们的研究结果凸显了整合异质标注和身体分割以提升模型性能的有效性,为3D医学影像中具备临床可靠性、数据高效的VLM指明了新方向。
英文摘要
General-purpose vision-language models (VLMs) have emerged as promising tools in radiology, offering zero-shot capabilities that mitigate the need for large labeled datasets. However, in high-stakes domains like diagnostic radiology, these models often lack the discriminative precision required for reliable clinical use. This challenge is compounded by the scarcity and heterogeneity of publicly available volumetric CT datasets, which vary widely in annotation formats and granularity. To address these limitations, we introduce Uniferum, a volumetric VLM that unifies diverse supervision signals, encoded in classification labels and segmentation masks, into a single training framework. By harmonizing three public 3D CT datasets with distinct annotations, Uniferum achieves state-of-the-art performance, improving AUROC on the CT-RATE benchmark by 7% compared to CLIP-based and conventional multi-label convolutional models. The model demonstrates robust out-of-distribution generalization, with observed evidence of unexpected zero-shot performance on the RAD-CHEST and INSPECT datasets. Our results highlight the effectiveness of integrating heterogeneous annotations and body segmentation to enhance model performance, setting a new direction for clinically reliable, data-efficient VLMs in 3D medical imaging.