arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07886cs.CVcs.AIcs.CL

视觉-语言 Grounding 作为双向概念对应

Vision-Language Grounding as Bidirectional Concept Correspondence

Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna

首次发表
浏览论文内容

中文总结 AI 辅助

本研究将视觉-语言 Grounding 建模为双向概念对应,提出 ConCor-1 模型统一相关任务,在长文本数据集和零样本 LVIS 上对应 F1 分别提升 48%、29%,性能优于基线。

中文摘要 AI 辅助

视觉-语言 Grounding 将语言与视觉内容关联起来,但现有大多数公式将 Grounding 简化为单向定位问题:给定预先指定的文本短语或类别名称,识别对应的图像区域。这种设置假设相关语言单元已被知晓,却忽略了 Grounded 通信中更基础的挑战:确定文本中哪些部分具有视觉指代性,以及它们如何对应图像中的实体。我们将 Grounding 公式化为图像-文本对的双向概念对应(bidirectional concept correspondence)。给定一张图像及其配对文本,目标是恢复所有具有视觉指代性的文本片段与实例级图像分割之间的对应关系,且不假设相关文本片段已被提供。该公式通过将文本分割、图像分割和跨模态对齐视为单一对应预测问题,统一了常见的 Grounding 任务,包括短语 Grounding、指代表达 Grounding 和开放词汇检测。为解决此任务,我们引入了基于预训练视觉-语言模型构建的 Grounding 模型 ConCor-1,它使用可学习的桥接令牌(bridge tokens)表示候选图像-文本对应关系,并为每个令牌预测文本掩码、图像掩码和对应存在分数。为训练和评估该任务,我们将多种 Grounding 和分割数据集转换为统一的对应格式。实验表明,ConCor-1 始终优于基线,在长文本数据集上的对应 F1 提升了 48%,在零样本 LVIS(其大类列表作为文本输入)上提升了 29%。

英文摘要

Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This setup assumes that the relevant linguistic unit is already known, overlooking a more basic challenge in grounded communication: determining which parts of the text are visually referential and how they correspond to entities in the image. We formulate grounding as $\textit{bidirectional concept correspondence}$ over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem. To address this task, we introduce $\textbf{ConCor-1}$, a grounding model built on top of a pretrained vision-language model. It uses learnable $\textit{bridge tokens}$ to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score. To train and evaluate this task, we convert diverse grounding and segmentation datasets into a unified correspondence format. Experiments show that $\textbf{ConCor-1}$ consistently outperforms baselines, improving correspondence F1 by 48% on the long-caption dataset and by 29% on zero-shot LVIS, where the large category list serves as the text input.

发表机构

  • University of Washington(华盛顿大学)
  • Allen Institute for AI(艾伦人工智能研究所)
  • FAIR at Meta(Meta FAIR实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑