arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15615cs.CV

用于检测引导的乳腺钼靶病变分类的区域基础视觉语言学习

Region-Grounded Vision-Language Learning for Detection-Guided Mammographic Lesion Classification

Zhengbo Zhou, Jiren Li, Dooman Arefan, Margarita Zuley, Shandong Wu

首次发表
浏览论文内容

中文总结 AI 辅助

针对乳腺钼靶病变分类,提出区域基础视觉语言学习方法,先通过区域-文本对比预训练对齐特征,引入多组件目标减轻偏差,再联合优化辅助病变检测头,实验表明该方法在多设置下性能优于相关方法。

中文摘要 AI 辅助

用对比目标训练的视觉语言模型在医学图像分析中显示出前景。然而,传统的全局图像-文本对齐不适用于乳腺钼靶检查,其中与诊断相关的病变在空间上是局部的,仅占图像的一小部分。在全图像级别学习表示时,对恶性评估至关重要的细微形态线索可能会被稀释。在这项工作中,我们提出了一种用于检测引导的乳腺钼靶病变分类的新型区域基础视觉语言学习方法。该方法模仿放射科医生的诊断范式。首先,一个区域-文本对比预训练阶段将病变特定特征与从放射学元数据派生的结构化临床描述符对齐。为了减轻低词汇量设置中的语义崩溃和背景偏差,我们引入了一个包含正对齐、细粒度语义硬负样本和背景抑制的多组件目标。其次,一个辅助病变检测头与对比分类联合优化,以保持空间敏感性并实现定位感知恶性分类。在两个独立数据集CBIS-DDSM和VinDr-Mammo上的大量实验表明,在域内、跨数据集和迁移学习设置下,我们的方法比相关方法具有更好的性能。

英文摘要

Vision-language models trained with contrastive objectives have shown promise in medical image analysis. However, conventional global image-text alignment is ill-suited for mammography, where diagnostically relevant lesions are spatially localized and occupy only a small fraction of the image. Subtle morphological cues critical for malignancy assessment can be diluted when representations are learned at the whole-image level. In this work, we propose a novel region-grounded vision-language learning method for detection-guided mammographic lesion classification. The method mirrors radiologists' diagnostic paradigm. First, a region-text contrastive pretraining stage aligns lesion-specific features with structured clinical descriptors derived from radiology metadata. To mitigate semantic collapse and background bias in low-vocabulary settings, we introduce a multi-component objective incorporating positive alignment, fine-grained semantic hard negatives, and background suppression. Second, an auxiliary lesion detection head is jointly optimized with contrastive classification to preserve spatial sensitivity and enable localization-aware malignancy classification. Extensive experiments on two independent datasets, CBIS-DDSM and VinDr-Mammo, show superior performance of our method compared to related methods under in-domain, cross-dataset, and transfer learning settings.

↑