arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26624cs.CV

文本到种子生成:通过将扩散模型重新用作文本引导的种子生成器实现免训练开放词汇种子语义分割

Text-to-seed generation: Training-free open-vocabulary seeded semantic segmentation via re-purposing diffusion as text-guided seed generator

Kumju Jo, Heesun Jung, Sungyong Baik

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出免训练框架T2S,利用Stable Diffusion生成文本引导的种子点,结合SAM实现开放词汇语义分割,在标准基准上取得优异性能。

中文摘要 AI 辅助

开放词汇语义分割(OVSS)旨在分割与任意文本查询对应的图像区域。尽管分割任何物体模型(SAM)是一种强大的分割基础模型,但它在OVSS上的独立性能仍然有限。因此,现有方法常使用SAM来优化其他模型预测的粗掩码,但当初始掩码不准确时,该策略不可靠。在本研究中,我们认为通过将SAM用作由准确物体点(即种子)而非不准确粗掩码引导的区域扩展模块,可实现更可靠的分割。受经典种子分割启发,我们将OVSS重新表述为文本引导的种子定位,随后基于种子的区域扩展。为实现这一思路,我们提出了Text-to-Seed(T2S),这是一种免训练框架,利用Stable Diffusion的文本-区域对应关系,为文本描述的目标类别生成基于注意力的种子点。随后,这些稀疏种子被用作SAM的点提示,以生成完整的物体掩码。无需特定任务训练或额外注释,T2S在标准OVSS基准上取得了优异性能,证明了将语义 grounding与种子驱动的空间分割相结合的有效性。

英文摘要

Open-vocabulary semantic segmentation (OVSS) aims to segment image regions corresponding to arbitrary text queries. Although the Segment Anything Model (SAM) is a powerful foundation model for segmentation, its standalone performance on OVSS remains limited. Existing methods therefore often use SAM to refine coarse masks predicted by other models, but this strategy is unreliable when the initial masks are inaccurate. In this work, we argue that more reliable segmentation can be achieved by exploiting SAM as a region expansion module guided by accurate object points (i.e., seeds) rather than inaccurate coarse masks. Inspired by classical seeded segmentation, we reformulate OVSS as text-guided seed localization followed by seed-based region expansion. To realize this idea, we propose Text-to-Seed (T2S), a training-free framework that leverages the text-to-region correspondence of Stable Diffusion to generate attention-based seed points for target categories described by text. These sparse seeds are then used as point prompts for SAM to produce full object masks. Without task-specific training or additional annotations, T2S achieves strong performance on standard OVSS benchmarks, demonstrating the effectiveness of combining semantic grounding with seed-driven spatial segmentation.

发表机构

  • Hanyang University(汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑