利用可分子集采样解决乳腺钼靶视觉语言模型中的“干草堆里找针”问题
Solving the Needle-in-a-Haystack Problem in Mammography Vision-Language Model with Differentiable Subset Sampling
浏览论文内容
中文总结 AI 辅助
本研究提出TopKSigLIP视觉语言模型,通过TopK-Patch模块与Sup-sigmoid损失解决乳腺钼靶VLM的性能问题,在多任务零样本评估中优于现有模型且定位效果更佳。
中文摘要 AI 辅助
将CLIP风格的视觉语言模型(VLM)预训练应用于乳腺钼靶的兴趣日益增长,但直接采用标准CLIP架构与训练目标的模型,在癌症、病灶类型及BI-RADS预测等临床重要任务中表现出有限的零样本性能。我们认为这种不佳表现源于忽略了乳腺钼靶数据的两个特征:(1)其高分辨率特性;(2)放射学报告的同质性,这在很大程度上由检查中阴性/良性病灶占主导所驱动。我们提出TopKSigLIP,一种旨在通过新颖架构与学习目标解决这两个限制的VLM。TopKSigLIP未将高分辨率乳腺钼靶图像下缩放以满足GPU内存约束,而是引入TopK-Patch模块,学习采样可能包含病灶的高分辨率图像块稀疏集合,规避了VLM训练中分辨率-批次大小的权衡;采样的图像块位置还可作为内置定位工具。为解决报告同质性问题,我们将错误排斥语义相似对的对比损失替换为Sup-sigmoid损失,该损失从SigLIP扩展sigmoid损失,使用结构化数据衍生的软标签。TopKSigLIP在密度评估、BI-RADS分类、病灶亚型及癌症预测任务的内部与外部基准上,零样本评估表现优于现有开源乳腺钼靶及通用医学VLM;尽管使用的视觉编码器显著更小、训练批次也小于基线,在线性探测下仍保持竞争力;TopK-Patch模块相较事后Grad-CAM还实现了更优的病灶定位。代码与权重已公开:this https URL。
英文摘要
There is growing interest in adopting CLIP-style vision--language model (VLM) pretraining for mammography. However, models that directly employ the standard CLIP architecture and training objective exhibit limited zero-shot performance in clinically important tasks such as cancer, finding-type, and BI-RADS predictions. We argue that this underwhelming performance is due to neglecting two characteristics of mammography data: (1) its high-res nature, and (2) homogeneity of radiology reports, largely driven by a predominance of negative/benign findings on examinations. We propose TopKSigLIP, a VLM designed to address these two limitations through a novel architecture and learning objectives. Instead of downscaling high-res mammography images to satisfy GPU memory constraints, TopKSigLIP introduces TopK-Patch module that learns to sample a sparse set of high-res patches likely to contain lesions, sidestepping the resolution--batch size tradeoff of VLM training. The sampled patch locations additionally serve as a built-in localization tool. To address report homogeneity, we replace the contrastive loss, which falsely repels semantically similar pairs, with a Sup-sigmoid loss. Sup-sigmoid loss extends the sigmoid loss from SigLIP with soft labels derived from structured data. TopKSigLIP outperforms existing open-source mammography and general medical VLMs on both internal and external benchmarks on density assessment, BI-RADS classification, finding subtyping, and cancer prediction under zero-shot evaluation. TopKSigLIP remains competitive under linear probing despite using a significantly smaller vision encoder and smaller training batches than baselines. The TopK-Patch module additionally achieves superior lesion localization over post-hoc Grad-CAM. Code and weights are made public:https://github.com/Youngseok0001/TopKSigLIP.
发表机构
- Emory University(埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。