arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboQuest:搜索、检查与测试的通用物理智能体

RoboQuest: Generalist Physical Agents that Search, Inspect and Test

Liu Renhang, Navonil Majumder, Tej Deep Pala, Soujanya Poria

arXiv 2610.10388首次发表:更新:

发表机构

Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RoboQuest是一个具身探索基准,要求智能体通过物理交互主动获取任务相关信息,涵盖搜索、检查和交互测试三种不确定性,评估显示最佳智能体成功率仅23%,并揭示过早停止探索等关键失败原因。

AI 中文摘要

多模态基础模型的最新进展使其成为能够执行一系列操作任务的通用物理智能体。然而,在陌生环境中成功运行可能要求智能体在观测中缺乏任务相关信息时通过交互主动寻求信息:它可能需要确定相关物体的位置、检查未观测到的属性,或发现不熟悉工具的效果。为此,我们引入RoboQuest,一个用于目标导向的具身探索的基准,其中智能体必须通过物理交互主动获取任务相关信息,利用所得证据调整后续行动,并自主决定何时提交任务完成。RoboQuest包含十个移动操作任务,聚焦于三种不确定性形式:搜索、基于操作的检查和交互式测试。我们通过一个通用的视觉运动接口评估了五个前沿多模态智能体,以及一个在我们发布的完整回合演示上微调的π_{0.5}策略。最佳智能体仅在23%的回合中成功,而微调策略几乎从未成功。对任务所基于的执行技能的孤立测试(提供隐藏信息)表明,智能体能够执行大多数所需动作,我们的失败分析仅将少数失败归因于执行。我们的失败分析进一步发现,智能体常常过早停止探索,因为它们在观察到任务完成所需证据之前就做出决策。我们还发现,智能体很少预防或修复其探索造成的干扰。此外,通过试错学习对大多数模型仍然困难。

英文摘要

Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a $π_{0.5}$ policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23\% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑