PANORAMA: 通过掩码提议选择的全景接地描述
PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
浏览论文内容
中文总结 AI 辅助
针对视觉-语言模型在图像描述中缺乏像素级接地的问题,提出全景接地描述任务,构建PanoCaps基准并设计PANORAMA模型,通过从短语条件掩码提议中选择实现高质量接地,达到最优性能。
中文摘要 AI 辅助
在现实世界中行动的智能系统需要对图像进行既全面又具有空间接地性的理解。当前的视觉-语言模型(VLM)能够生成流畅且详细的图像描述,但将其可靠地与图像像素关联起来仍然具有挑战性。现有的将密集描述与像素级接地相结合的方法,往往会产生不完整的描述或不准确的分割掩码。我们通过全景接地描述这一任务来研究这个问题,该任务要求VLM描述前景物体和背景区域,同时将每个指代短语与像素级掩码进行接地。我们做出了三项贡献。首先,我们引入了PanoCaps,这是一个从全景分割数据集构建的人工标注基准。它提供了具有近乎完整像素覆盖的密集描述和实体级别的图像-文本对齐,支持训练和评估。我们进一步提出了一种短语-掩码匹配协议和一种广义全景质量(gPQ)指标,该指标联合评估文本和掩码的一致性。其次,我们将短语接地表述为从短语条件下的掩码提议池中进行选择,并引入了PANORAMA,这是一种VLM,它使预训练的分割器以上下文化的短语表示为条件来获取候选掩码,并学习选择与每个短语对应的掩码。将此接口与描述生成联合训练,使PANORAMA能够生成高质量的掩码,同时允许每个短语指代单个区域或多个实例。第三,PANORAMA在PanoCaps上实现了最佳的整体接地性能,并在多个像素级接地任务上达到或超过了专门模型。实验表明,我们的方法能够生成精确的实体级分割,同时保持详细且与掩码一致的描述。代码、数据和模型可在以下网址获取:https URL。
英文摘要
Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense captioning with pixel-level grounding often produce either incomplete descriptions or inaccurate segmentation masks. We study this problem through panoptic grounded captioning, a task that requires a VLM to describe both foreground objects and background regions while grounding each referring phrase with pixel-level masks. We make three contributions. First, we introduce PanoCaps, a human-annotated benchmark constructed from panoptic segmentation datasets. It provides dense captions with near-complete pixel coverage and image-text alignments at the entity level, supporting both training and evaluation. We further propose a phrase-mask matching protocol and a generalized Panoptic Quality (gPQ) metric that jointly evaluates textual and mask agreement. Second, we formulate phrase grounding as selection from a phrase-conditioned pool of mask proposals and introduce PANORAMA, a VLM that conditions a pretrained segmenter on contextualized phrase representations to obtain candidate masks and learns to select those corresponding to each phrase. Training this interface jointly with caption generation enables PANORAMA to produce high-quality masks while allowing each phrase to refer to a single region or multiple instances. Third, PANORAMA achieves the best overall grounding on PanoCaps and matches or exceeds specialized models across several pixel-level grounding tasks. Experiments show that our method produces precise entity-level segmentations while maintaining detailed, mask-consistent captions. Code, data and models are available at https://www.di.ens.fr/willow/research/panorama/.
发表机构
- Inria, École normale supérieure, CNRS, PSL Research University(法国国家信息与自动化研究所,巴黎高等师范学院,法国国家科学研究中心,巴黎文理研究大学)
- Czech Institute of Informatics, Robotics and Cybernetics, Czech Technical University in Prague(捷克信息学、机器人与控制论研究所,布拉格捷克理工大学)
机构由 AI 辅助整理,请以论文原文为准。