arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02284cs.CV

EOVSAM:基于SAM 3的单遍高效开放词汇分割

EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass

Haomin Peng, Yongkang Li, Zhaoxiang Liu, Xiaojie Jin, Shiguo Lian, Yunchao Wei, Xinggang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

EOVSAM是适配SAM 3的高效开放词汇分割框架,通过单遍预测、注意力聚合策略提升精度,推理速度最高提338倍,在多基准上兼具高精度与快速度优势。

中文摘要 AI 辅助

开放词汇分割是指从任意文本描述中识别并分割出对应物体的任务。SAM 3支持名词短语引导的分割,通过穷尽遍历词汇表实现了颇具竞争力的开放词汇性能,但随着目标类别规模扩大,其计算开销会变得难以承受。本文提出了一种基于SAM 3的高效开放词汇分割框架EOVSAM,该框架将SAM 3适配为单遍预测模式,移除了提示条件以将SAM 3转化为高效掩码生成器,并引入了全新的注意力聚合策略来端到端优化开放词汇分类。这种设计避免了现有方法常用的多阶段流水线和后处理启发式方法,同时缓解了直接优化分类时可能出现的闭集塌陷问题。在所有评估数据集上,EOVSAM均持续提升了原始SAM 3的分割精度,推理速度最高提升了338倍;此外,EOVSAM在较低分辨率下仍能保持高精度,推理速度优势更为显著。在标准语义和全景分割基准上的实验表明,EOVSAM兼具与现有开放词汇分割模型相当或更优的精度,同时拥有显著的速度优势。代码和模型可在该https链接获取。

英文摘要

Open-vocabulary segmentation identifies and segments objects from arbitrary textual descriptions. SAM 3 supports noun-phrase-guided segmentation and achieves competitive open-vocabulary performance through exhaustive vocabulary traversal, yet suffers from prohibitive computational overhead as target categories scale. In this paper, we propose an Efficient Open-Vocabulary segmentation framework with SAM 3 (EOVSAM), which adapts SAM 3 for single-pass prediction. EOVSAM removes prompt conditioning to turn SAM 3 into an efficient mask generator and introduces a new Attentional Aggregation strategy to optimize open-vocabulary classification end-to-end. This formulation avoids the multi-stage pipelines and post-processing heuristics commonly used by existing methods, while mitigating the closed-set collapse that can arise when classification is optimized directly. EOVSAM consistently improves segmentation accuracy over vanilla SAM 3 on all evaluated datasets and accelerates inference by up to 338$\times$. Furthermore, EOVSAM maintains high accuracy at lower resolutions while achieving even more remarkable inference speeds. Experiments on standard semantic and panoptic segmentation benchmarks show that EOVSAM combines competitive or state-of-the-art accuracy with a substantial speed advantage over existing open-vocabulary segmentation models. Code and models are available at https://github.com/hustvl/EOVSAM.

↑