arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09478cs.CV

InstanceBench:诊断指称表达分割中的指称推理与目标同一性

InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation

发表机构悉尼大学 · ATLAS研究所 · 上海人工智能实验室
查看机构详情
  • The University of Sydney(悉尼大学)
  • The ATLAS Institute(ATLAS研究所)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yuchen Li, Shaoyang Zhou, Yiran Wang, Ruiyi Deng, Haoyu Wang, Ziru Wei, Zhen Zhao, Luping Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

InstanceBench是一个以实例为中心的诊断基准,通过指称逻辑分类法和身份感知指标,系统评估RES模型的目标选择能力,发现目标选择是主要瓶颈,并支持测量-诊断-改进循环。

中文摘要 AI 辅助

指称表达分割(RES)将自然语言描述与像素级对象掩码相关联。然而,标准评估对实例级指称推理的洞察有限:它不能系统地区分指称逻辑,不能测试跨有效接地路径的目标保持性,也不能将目标选择与掩码生成错误分开。我们引入了InstanceBench,一个以实例为中心的诊断基准,包含6,194张图像、9,264个目标实例和25,077条人工验证的表达。每个以目标为中心的表达式集(TCES)固定图像和目标掩码,同时将最小表达式与使用另一种有效线索或接地路径的相同目标变体配对。一个紧凑的指称逻辑分类法涵盖直接目标证据、同类选择、关系和组合接地以及排除,而逻辑关键的构造抑制了更简单的捷径。身份感知指标衡量目标保持和集合级成功,同时将选择与掩码生成错误分开。在来自18个模型家族的22个原生掩码RES检查点中,最强的检查点达到67.1%的mIoU,但只有59.6%的All@0.7。受控干预确认了语言敏感性,而失败分解确定目标选择而非掩码解码是主要瓶颈。在受控训练子集上,匹配的监督提高了身份感知性能,表明被诊断的能力响应于有针对性的监督。总的来说,InstanceBench支持一个测量-诊断-改进循环:测量跨接地路径的目标一致性,定位失败源,并评估有针对性的干预。

英文摘要

Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separate target-selection from mask-generation errors. We introduce InstanceBench, an instance-centered diagnostic benchmark comprising 6,194 images, 9,264 target instances, and 25,077 human-verified expressions. Each target-centric expression set (TCES) fixes the image and target mask while pairing a minimal expression with a same-target variant that uses another valid cue or grounding path. A compact referential-logic taxonomy spans direct target evidence, same-class selection, relational and compositional grounding, and exclusion, while logic-critical construction suppresses simpler shortcuts. Identity-aware metrics measure target retention and set-level success while separating selection from mask-generation errors. Across 22 native-mask RES checkpoints from 18 model families, the strongest checkpoint reaches 67.1% mIoU but only 59.6% All@0.7. Controlled interventions confirm language sensitivity, while failure decomposition identifies target selection rather than mask decoding as the main bottleneck. On a controlled training subset, matched supervision improves identity-aware performance, showing that the diagnosed capability responds to targeted supervision. Collectively, InstanceBench supports a measure-diagnose-improve cycle: measuring target consistency across grounding paths, localizing failure sources, and evaluating targeted interventions.

↑