arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MARS-CLIP:多分辨率与注意力精炼的零样本图像分割

MARS-CLIP: Multi-Resolution and Attention Refined Zero-Shot Image Segmentation

Nagito Saito, Shintaro Ito, Koichi Ito, Takafumi Aoki

arXiv 2609.08283首次发表:更新:

发表机构

Graduate School of Information Sciences, Tohoku University(东北大学信息科学研究生院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MARS-CLIP框架,通过多分辨率特征融合与注意力精炼机制,解决CLIP在零样本语义分割中分辨率低和结构信息丢失的问题,在六个数据集上超越现有方法。

AI 中文摘要

对比语言-图像预训练(CLIP)在零样本迁移中展现了令人印象深刻的能力,但由于低空间分辨率和结构信息的丢失,在处理密集预测任务时常常遇到困难。为了解决这些局限性,我们提出了MARS-CLIP(多分辨率与注意力精炼的CLIP分割),一种用于零样本语义分割的新颖框架。我们的方法引入了两个关键策略:(i)一个多分辨率特征提取模块,将局部细粒度特征与全局上下文融合,以克服输入分辨率限制;(ii)一种注意力精炼机制,将中间层的空间和颜色偏差注入到最终的自注意力块中,以准确恢复对象边界。在六个公共数据集上进行的一系列实验表明,MARS-CLIP显著优于最先进的方法。

英文摘要

Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in zero-shot transfer but often struggles with dense prediction tasks due to low spatial resolution and the loss of structural information. To address these limitations, we propose MARS-CLIP (Multi-resolution and Attention Refined Segmentation for CLIP), a novel framework for zero-shot semantic segmentation. Our approach introduces two key strategies: (i) a multi-resolution feature extraction module that fuses local fine-grained features with global context to overcome input resolution constraints, and (ii) an attention refinement mechanism that injects spatial and color biases from intermediate layers into the final self-attention block to accurately restore object boundaries. A set of experiments on six public datasets demonstrates that MARS-CLIP significantly outperforms state-of-the-art methods.

CommentsAccepted to IEEE International Conference on Image Processing (ICIP) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑