arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习多模态伪标签以实现鲁棒的开放词汇实例和全景分割

Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation

Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang

arXiv 2608.11681首次发表:更新:

发表机构

Seoul National University of Science and Technology; Chung-Ang University(首尔科技大学; 中央大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对开放词汇实例分割与开放集全景分割的痛点,提出多模态框架,结合预训练视觉语言模型生成伪标签并优化训练目标,在COCO数据集上取得优于现有SOTA的性能。

AI 中文摘要

本研究针对开放词汇实例分割(OVIS)和开放集全景分割(OSPS)面临的挑战展开,这两项任务旨在无需详尽人工标注的前提下,识别预定义及未见过的物体类别。现有方法常受限于伪掩码噪声、视觉-文本对齐不足,以及难以处理同义词或词汇外(OOV)词汇的问题。为克服这些挑战,我们提出一种多模态框架,利用预训练视觉-语言模型生成自动伪标签、基于CLIP的同义词过滤,以及基于GPT的标题重构。在目标词汇辅助的伪标签设置中,该框架首先使用Grounded SAM、LLaVA和CLIP构建伪分割掩码、描述性标题及语义对齐的同义词集,无需人工标注即可提供多模态监督。随后,我们通过三个互补的训练目标增强视觉-文本对齐:包含视觉对齐同义词的扩展对齐损失、语义一致性损失,以及生成式标题重构损失。在COCO数据集上开展的大量实验表明,所提方法在该协议下始终优于现有最先进方法,在OVIS和OSPS基准上均实现了显著提升。

英文摘要

This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to recognize both predefined and unseen object categories without exhaustive human annotations. Existing methods often suffer from noisy pseudo-masks, limited visual-textual grounding, and difficulty handling synonyms or out-of-vocabulary (OOV) words. To overcome these challenges, we propose a multimodal framework that leverages pre-trained vision-language models for automatic pseudo-label generation, CLIP-guided synonym filtering, and GPT-based caption reconstruction. In our target-vocabulary-assisted pseudo-labeling setting, the framework first constructs pseudo segmentation masks, descriptive captions, and semantically aligned synonym sets using Grounded SAM, LLaVA, and CLIP, providing multimodal supervision without manual annotation. We then enhance visual-textual alignment through three complementary training objectives: an extended grounding loss that incorporates visually grounded synonyms, a semantic consistency loss, and a generative caption reconstruction loss. Extensive experiments on the COCO dataset demonstrate that the proposed method consistently outperforms previous state-of-the-art approaches under this protocol, achieving substantial improvements on both OVIS and OSPS benchmarks.

Comments14 pages

Journal refNeurocomputing, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑