arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VespaSeg:一种用于指代表达分割的资源感知“先定位后分割”流水线

VespaSeg: A Resource-Aware Ground-then-Segment Pipeline for Referring Expression Segmentation

Savindu Dilshan Wickramasinghe

arXiv 2608.01077首次发表:更新:

发表机构

University of Moratuwa(莫拉图瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对指代表达分割中整体式模型部署成本高的问题,提出模块化流水线VespaSeg,采用轻量型视觉-语言模型定位与MobileSAM分割,经适配后在RefCOCO数据集上实现高精度,且内存占用低、推理速度快。

AI 中文摘要

指代表达分割需要以语言为条件的定位以及像素级精确的掩码,但整体式模型的部署成本较高。本文提出VespaSeg,这是一种模块化流水线,它用轻量型视觉-语言模型对文本查询进行定位,再用MobileSAM将预测框转换为掩码。我们研究了Florence-2-base、Florence-2-large和Moondream2定位器,并对定位和分割阶段进行针对性适配。在包含3811个被引用对象记录的首个表达式的特定仓库RefCOCO验证协议下,适配后的Florence-2-base流水线获得73.64的平均交并比(mIoU)和IoU为0.5时84.60的精度。在NVIDIA RTX 6000 Ada GPU上,它每秒可处理22.8个缓存图像查询,平均分配GPU内存为2.20 GB。在500个查询的匹配对比中,Florence-2-base的mIoU为73.73,Florence-2-large为72.82,同时基础模型速度快1.70倍,分配内存少1.17 GB。 ablation实验显示,真实框适配将MobileSAM的mIoU从82.22提升至86.61,将Florence-2输出令牌预算从64减少至32可保持精度。这些结果为轻量型模块化定位与分割提供了支撑,同时也凸显了在完整标准RefCOCO表达式划分和部署硬件上进行评估的必要性。

英文摘要

Referring expression segmentation requires language conditioned localization and pixel-accurate masks, but monolithic models can be costly to deploy. We present VespaSeg, a modular pipeline that grounds a text query with a compact vision-language model and converts the predicted box to a mask with MobileSAM. We study Florence-2-base, Florence-2-large, and Moondream2 grounders together with targeted adaptation of the grounding and segmentation stages. Under a repository-specific RefCOCO validation protocol containing the first expression for each of 3,811 referenced-object records, the adapted Florence-2-base pipeline obtains 73.64 mean intersection over union (mIoU) and 84.60 precision at IoU 0.5. On an NVIDIA RTX 6000 Ada GPU it processes 22.8 cached-image queries per second with 2.20 GB mean allocated GPU memory. A matched 500-query comparison gives 73.73 mIoU for Florence-2-base and 72.82 for Florence-2-large, while the base model is 1.70 times faster and uses 1.17 GB less allocated memory. Ablations show that ground-truth-box adaptation raises MobileSAM mIoU from 82.22 to 86.61 and that reducing the Florence-2 output-token budget from 64 to 32 preserves accuracy. These results support compact, modular grounding and segmentation, while also exposing the need for evaluation on the complete standard RefCOCO expression splits and deployment hardware.

Comments5 pages, 3 figures. Code and result artifacts: https://github.com/Savidilsh/VespaSeg

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑