arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当视觉不再足够:评估大视觉语言模型(LVLMs)中的交互式视觉定位

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

Zhengxiang Wang, Owen Rambow

arXiv 2608.23978首次发表:更新:

发表机构

Stony Brook University(石溪大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对大视觉语言模型(LVLMs)提出交互式视觉定位的受控评估框架,发现当前LVLMs表现低于人类基线,主动提问式定位困难且校准不佳,该任务仍具挑战性。

AI 中文摘要

视觉定位通常被评估为从信息丰富的指代表达到视觉目标的一次性映射。这种表述忽略了现实世界指称的核心属性:目标信息往往不完整、模糊,且通过交互建立。我们为大视觉语言模型(LVLMs)引入了一个受控的交互式视觉定位评估框架,改变预先提供的目标信息量以及必须通过对话获取的信息量。在四个人类标注的视觉上下文和四种交互协议下,当前LVLMs的表现显著低于任务级人类基线。当后续问题完善或修正初始目标描述时,交互会有帮助;当未提供初始描述且必须通过提问获取目标信息时,性能最低,表明主动提问驱动的定位仍然困难。LVLMs的校准也较差,常报告超出其经验准确率的置信度。后续研究在不同描述来源(人类与AI)、推理力度、重复交互、描述提供者和视觉上下文中均证实了这些模式。总体而言,交互式视觉定位仍是重要挑战,需要视觉匹配、信息获取与综合能力。

英文摘要

Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: initial referring expressions are often incomplete or ambiguous, requiring participants to establish shared understanding through interaction. We introduce a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), varying how much target information is provided upfront and how much must be acquired through dialogue. Across four human-grounded visual contexts and four interaction protocols, current LVLMs perform significantly below task-level human baselines. Interaction can help when follow-up questions refine or repair an initial target description. Performance is lowest when no initial description is provided and target information must be acquired through questions, indicating that proactive question-driven grounding remains difficult. LVLMs are also poorly calibrated, often reporting confidence that exceeds their empirical accuracy. Follow-up studies confirm these patterns across varied description sources (human versus AI), reasoning efforts, repeated interactions, description providers, and visual contexts. Overall, interactive visual grounding remains challenging, requiring visual matching, information seeking and synthesis.

CommentsEMNLP 2026 Main

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑