发表机构
Graduate School of Information Sciences, Tohoku University(东北大学信息科学研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MARS-CLIP框架,通过多分辨率特征融合与注意力精炼机制,解决CLIP在零样本语义分割中分辨率低和结构信息丢失的问题,在六个数据集上超越现有方法。
AI 中文摘要
对比语言-图像预训练(CLIP)在零样本迁移中展现了令人印象深刻的能力,但由于低空间分辨率和结构信息的丢失,在处理密集预测任务时常常遇到困难。为了解决这些局限性,我们提出了MARS-CLIP(多分辨率与注意力精炼的CLIP分割),一种用于零样本语义分割的新颖框架。我们的方法引入了两个关键策略:(i)一个多分辨率特征提取模块,将局部细粒度特征与全局上下文融合,以克服输入分辨率限制;(ii)一种注意力精炼机制,将中间层的空间和颜色偏差注入到最终的自注意力块中,以准确恢复对象边界。在六个公共数据集上进行的一系列实验表明,MARS-CLIP显著优于最先进的方法。
英文摘要
Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in zero-shot transfer but often struggles with dense prediction tasks due to low spatial resolution and the loss of structural information. To address these limitations, we propose MARS-CLIP (Multi-resolution and Attention Refined Segmentation for CLIP), a novel framework for zero-shot semantic segmentation. Our approach introduces two key strategies: (i) a multi-resolution feature extraction module that fuses local fine-grained features with global context to overcome input resolution constraints, and (ii) an attention refinement mechanism that injects spatial and color biases from intermediate layers into the final self-attention block to accurately restore object boundaries. A set of experiments on six public datasets demonstrates that MARS-CLIP significantly outperforms state-of-the-art methods.
CommentsAccepted to IEEE International Conference on Image Processing (ICIP) 2026