arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30621cs.CVcs.AI

用于指称图像分割与指称的高成本效益主动学习

Cost-efficient Active Learning for Referring Image Segmentation and Grounding

  • Meissa
  • KAIST(韩国科学技术院)
  • GIST(光州科学技术院)
  • POSTECH(浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Junbeom Hong, Seonghoon Yu, Hyung Rok Jung, Sundong Kim, Jeany Son

AI总结:

该研究针对视觉指称的标注瓶颈,提出结合基础模型生成辅助区域-文本对与指称区域模糊度获取函数的主动学习框架,在RIS、REC基准及用户研究中均实现更优性能与更快标注速度。

AI中文摘要:

收集自然语言指称表达式以及区域标注(如掩码或边界框)是视觉指称(VG)领域的主要瓶颈,因为标注者必须编写能够将目标区域与视觉相似区域区分开的描述。我们针对仅存在原始图像、无对应文本的现实场景,研究视觉指称下的主动学习(AL)以解决该问题。由于不存在真实文本,样本选择需评估哪些图像包含需要区分性指称表达式的模糊区域。为解决此问题,我们使用基础模型生成辅助区域-文本对,并引入新的获取函数“指称区域模糊度(Referred Region Ambiguity)”,用于衡量模型的置信度是集中于单个区域还是分散在多个候选区域。该函数使我们的方法能够优先选择具有强跨区域竞争的图像,这类图像因视觉模糊性而更具信息价值。我们还设计了指称表达式标注界面,帮助标注者通过少量点击快速聚焦于编写区分性语言。在RIS和REC基准上的实验表明,我们的主动学习框架始终优于多个主动学习基线,而用户研究显示我们的方法可使描述标注速度提升最高达1.6倍。

英文摘要:

Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target regions from visually similar ones. We tackle this by formulating active learning (AL) for VG under the realistic setting where only raw images are available without accompanying text. Since ground-truth text is unavailable, sample selection must estimate which images contain ambiguous regions that would require discriminative referring expressions. To address this, we generate auxiliary region-text pairs using foundation models, and introduce Referred Region Ambiguity, a new acquisition function that measures whether the model's confidence collapses onto a single region or disperses across multiple candidates. It allows our method to prioritize images with strong cross-region competition, which are more informative due to their visual ambiguity. We also design a referring-expression annotation interface that helps annotators quickly focus on writing discriminative language with a few clicks. Experiments on RIS and REC benchmarks show that our AL framework consistently outperforms several AL baselines, while a user study shows up to 1.6X faster description labeling of ours.

补充信息

↑