arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Diffuse2Seg:扩散模型可在无监督条件下分割任意目标

Diffuse2Seg: Diffusion Models Can Segment Anything Without Supervision

Christoph Hümmer, Joachim Sicking, Fabian Hüger, Hanno Gottschalk

arXiv 2609.06491首次发表:更新:

发表机构

Institute of Mathematics, Technical University of Berlin; CARIAD SE, Volkswagen Group(柏林工业大学数学学院; 大众集团旗下CARIAD SE)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Diffuse2Seg 利用扩散模型自注意力表征传播点提示生成多粒度掩码,无需监督即可分割任意目标,在多个基准上超越现有标签生成器并提升下游分割性能。

AI 中文摘要

开放世界实体分割旨在跨领域、多粒度(从部件到整体目标)预测任意目标的掩码。在此设定下,SAM 树立了强大标杆:其在包含 1100 万张图像和超过 10 亿个精心标注掩码的 SA-1B 上训练,实现了显著的零样本性能。然而,收集此类标注成本高昂且耗时,限制了该方法的可扩展性。文本到图像扩散模型为此提供了解决途径。它们的中间特征在感知任务间迁移良好,且由于目标结构在模型将噪声样本去噪为受文本提示条件约束的图像过程中涌现,该结构已编码于这些表征中。因此,这些表征可用于开放世界实体分割,无需重新训练或监督。基于此观察,我们提出 Diffuse2Seg,它通过以边缘保持方式将点提示网格传播至自注意力表征,重新利用生成式扩散模型进行自动掩码生成。Diffuse2Seg 生成多粒度实例掩码,在五个域上 AR_1000 指标超越先前最先进的标签生成器 4.3-7.1 个百分点。在这些生成掩码上训练实例分割模型,在“things”和“stuff+things”数据集上将无检测器开放世界分割分别提升 7.4 和 7.7 个百分点,并在“stuff+things”上 AR_1000 超越基于检测器的 UnSAM 2.1 个百分点。最后,我们表明在 Diffuse2Seg 标签上训练的模型为半监督学习提供强初始化,仅用 5k 张标注图像即超越其全监督对应模型。

英文摘要

Open-world entity segmentation aims to predict masks for arbitrary objects across domains and at multiple granularities, from parts to whole objects. In this setting, SAM sets a strong standard: trained on SA-1B, comprising 11M images and over 1B carefully annotated masks, it achieves remarkable zero-shot performance. Collecting such labels is expensive and time-consuming, however, which limits how far this recipe can scale. Text-to-image diffusion models offer a way around this. Their intermediate features transfer well across perception tasks, and since object structure emerges as the model denoises a noise sample into an image conditioned on a text prompt, that structure is already encoded in these representations. They can therefore be exploited for open-world entity segmentation without retraining or supervision. Building on this observation, we introduce Diffuse2Seg, which repurposes generative diffusion models for automatic mask generation by propagating a grid of point prompts through their self-attention representations in an edge-preserving manner. Diffuse2Seg produces multi-granular instance masks and outperforms prior state-of-the-art label generators by 4.3-7.1 p.p. in AR_1000 across five domains. Training an instance segmentation model on these generated masks advances detector-free open-world segmentation by 7.4 and 7.7 p.p. on "things" and "stuff+things" datasets and surpasses the detector-based UnSAM on "stuff+things" by 2.1 p.p. in AR_1000. Finally, we show that a model trained on Diffuse2Seg labels provides a strong initialization for semi-supervised learning, outperforming its fully supervised counterpart with already 5k labeled images.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑