arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12454cs.CVcs.AIcs.LG

弥合视觉基础模型先验与CLIP:医学图像中空间感知的少样本异常检测

Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images

Juzheng Miao, Yuchen Yuan, Cheng Chen, Pheng-Ann Heng

首次发表
浏览论文内容

中文总结 AI 辅助

提出Spatial-FAD,融合DINO空间先验与CLIP语义,通过适配器、滑动窗口和原型记忆,在少样本医学图像异常检测中显著提升病灶分割性能。

中文摘要 AI 辅助

视觉语言模型(如CLIP)通过强大的图像-文本语义对齐,实现了有效的少样本医学异常检测(AD)。然而,其全局对比预训练缺乏显式的空间监督,限制了精确的病灶定位。相比之下,视觉基础模型(VFM)如DINO,通过自蒸馏和局部到全局一致性学习空间连贯的块级表示,能更好地捕捉细粒度的解剖结构。利用这种互补性,我们提出了Spatial-FAD,一种空间感知的少样本医学异常检测框架,通过结合VFM空间先验与CLIP语义来改进病灶定位。具体而言,我们引入了一个VFM增强适配器,将源自DINO的结构亲和先验注入CLIP特征中。这种结构引导的细化鼓励视觉嵌入更好地贴合病灶边界,同时保持语义对齐。为了解决分块化导致的空间细节丢失以及CLIP输入分辨率有限的问题,我们采用了滑动窗口聚合策略。该策略生成高分辨率、空间密集的嵌入,进一步增强定位的粒度。此外,我们引入了一种原型增强的支持记忆方案,以高效利用少样本支持集。该模块存储正常和异常模式的紧凑原型,通过融合块到原型和图像-文本相似性,在降低内存成本的同时提升性能。在三个基准数据集(包括肝脏CT、视网膜OCT和脑MRI)上的大量实验表明,Spatial-FAD显著优于最先进的方法,尤其在病灶分割方面。值得注意的是,在4-shot场景中,我们的方法在Dice分数上平均提升了超过11.4%,在AUC上提升了1.8%。代码可在以下网址获取:this https URL。

英文摘要

Vision-Language Models such as CLIP enable effective few-shot medical anomaly detection (AD) via strong image-text semantic alignment. However, their globally contrastive pretraining lacks explicit spatial supervision, limiting precise lesion localization. In contrast, Vision Foundation Models (VFMs) such as DINO learn spatially coherent patch representations via self-distillation and local-to-global consistency, better capturing fine-grained anatomical structures. Leveraging this complementarity, we propose Spatial-FAD, a spatial-aware few-shot medical AD framework that improves lesion localization by combining VFM spatial priors with CLIP semantics. Specifically, we introduce a VFM-enhanced adapter that injects a structural affinity prior derived from DINO into CLIP features. This structure-guided refinement encourages visual embeddings to better adhere to lesion boundaries while maintaining semantic alignment. To address the loss of spatial detail from patchification and the limited input resolution of CLIP, we adopt a sliding-window aggregation strategy. This generates high-resolution, spatially dense embeddings to further enhance localization granularity. Moreover, we introduce a prototype-enhanced support memory scheme to efficiently exploit the few-shot support set. This module stores compact prototypes for normal and abnormal patterns, reducing memory costs while boosting performance by fusing patch-to-prototype and image-text similarities. Extensive experiments on three benchmark datasets, including Liver CT, Retinal OCT, and Brain MRI, demonstrate that Spatial-FAD significantly outperforms state-of-the-art methods, especially in lesion segmentation. Notably, in the 4-shot scenario, our method achieves an average improvement of over 11.4% in Dice score and 1.8% in AUC. Code is available at: https://github.com/JuzhengMiao/Spatial-FAD.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)
  • The University of Hong Kong(香港大学)
  • Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学医学智能与扩展现实研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑