arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Pro-Bench:面向真实异构环境的提示鲁棒开放词汇视觉定位基准

Pro-Bench: Prompt-Robust Open-Vocabulary Visual Grounding Across Real-World Heterogeneous Environments

Linus Nwankwo, Muslim Alaran, Christian Rauch, Stanley Chukwuebuka Obilikpa, Elmar Rueckert

arXiv 2609.27076首次发表:更新:

发表机构

Montanuniversität Leoben; Middlesex University(莱奥本矿业大学; 密德萨斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Pro-Bench基准,包含多领域RGB帧和丰富查询,系统评估16种开放词汇模型在真实环境中的提示鲁棒性,发现其强依赖架构,并揭示聚合mAP掩盖的恢复差异。

AI 中文摘要

开放词汇视觉定位使机器人能够根据自然语言查询定位与任务相关的实体,而无需依赖预定义的感知分类体系。然而,现有基准大多依赖短类别标签和网络抓取的图像,尚不清楚开放词汇模型在真实部署中能否稳健地处理多样查询和视觉条件。我们提出Pro-Bench,一个面向异构真实环境的提示条件开放词汇视觉定位基准。Pro-Bench包含来自独立机器人领域(地下、工业、室内、室外、城市)的13000多帧RGB图像,具有74500个手动实例标注和515个目标查询,涵盖类别、属性、关系、可供性、状态、部分-整体、否定和组合语义。我们在严格的零样本推理下对16种开放词汇模型配置进行了基准测试,测量了跨IoU阈值的定位精度、端到端推理延迟、提示引起的性能变化以及目标恢复一致性。结果表明,提示鲁棒性强烈依赖于架构。大多数模型配置(16种中的10种)在短类别标签下表现最佳,而自由形式查询仅对一种配置产生最高精度。此外,相似的聚合mAP可能掩盖不同改写下目标恢复一致性的显著差异。Pro-Bench能够系统评估这些差距,并支持提示鲁棒的视觉定位。Pro-Bench:https://pro-bench.github.io/。

英文摘要

Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perceptual taxonomies. However, existing benchmarks largely rely on short category labels and web-scraped imagery, leaving it unclear whether open-vocabulary models can robustly ground diverse queries and visual conditions under real deployments. We introduce \textbf{Pro-Bench}, a prompt-conditioned benchmark for open-vocabulary visual grounding in heterogeneous, real-world environments. Pro-Bench includes $13k+$ RGB frames from independent robotic domains (subterranean, industrial, indoor, outdoor, urban), with $74.5k$ manual instance annotations and $515$ target queries covering categorical, attributive, relational, affordance, state, part-whole, negative, and compositional semantics. We benchmarked $16$ open-vocabulary model configurations in strict zero-shot inference, measuring localisation accuracy across IoU thresholds, end-to-end inference latency, prompt-induced performance variation, and target recovery consistency. Our results show that prompt-robustness is strongly architecture-dependent. Most model configurations ($10/16$) perform best with short category labels, whereas free-form queries yield the highest accuracy for only one. Moreover, similar aggregate mAP can conceal substantial differences in consistent target recovery across reformulations. Pro-Bench enables systematic evaluation of these gaps and supports prompt-robust visual grounding. Pro-Bench: https://pro-bench.github.io/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑