JRDB-AVR:真实世界环境中具身智能体的主动视觉推理基准
JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
- Monash University(莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对具身智能体在真实环境中因视野受限而缺乏主动证据获取的问题,提出JRDB-AVR基准和JRDB-AVR-Agent方法,通过图世界模型和求解实现主动视觉推理,实验显示当前模型存在答案与证据准确率差距,强调主动证据评估的必要性。
AI中文摘要:
在复杂的具身视觉推理场景中,智能体通常只有有限的视野,而回答问题所需的证据可能分布在时间、视角和交互对象上。因此,模型可能在未观察到相关对象、时间或视角的情况下给出看似合理的答案。当前的视觉推理基准主要评估被动观察和最终答案,忽视了需要主动推理和证据获取的场景。我们提出了JRDB-AVR,这是一个从现有真实世界JRDB机器人数据中通过结构化问题生成引擎构建的基准,将这一差距转化为明确的评估:具身智能体系统接收视觉推理问题,按时间戳和视角请求有界观察,并根据最终答案和支撑答案的基于视觉证据进行评估。该基准包含多个真实世界环境中多样的问题,涉及时间搜索、视角选择以及面向人的组合推理。我们还提出了JRDB-AVR-Agent,一种参考主动推理智能体方法,该方法维护一个显式的基于观察的图世界模型,并通过求解来回答问题。实验揭示了当前基线在答案准确性和证据准确性之间存在显著差距,表明当前的视觉语言模型可能产生无证据支持的正确答案,并且主动的、基于证据的评估对于具身视觉推理是必要的。代码和基准可在以下网址获取:https://github.com/...(此处为URL)
英文摘要:
In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that supports it. Current visual reasoning benchmarks largely evaluate passive observations and final answers, overlooking settings that require active reasoning and evidence acquisition. We introduce JRDB-AVR, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap into an explicit evaluation: an embodied agentic system receives a visual reasoning question, requests bounded observations by timestamp and viewing angle, and is evaluated on both the final answer and the grounded visual evidence supporting it. The benchmark contains diverse questions over multiple real-world environments involving temporal search, viewpoint selection, and human-oriented compositional reasoning. We also introduce JRDB-AVR-Agent, a reference active reasoning agentic method that maintains an explicit observation-grounded graph-based world model and answers through solving. Experiments reveal a substantial gap between answer accuracy and evidence accuracy in current baselines, showing that current VLMs can produce unsupported correct answers and that active evidence-aware evaluation is necessary for embodied visual reasoning. Code and benchmark are available at https://github.com/ControlNet/JRDB-AVR.