RMMBench:机器人移动操作的综合基准
RMMBench: A Comprehensive Benchmark for Robotic Mobile Manipulation
浏览论文内容
中文总结 AI 辅助
针对现有基准缺乏全面评估方法的问题,提出RMMBench基准,集成高低层具身任务,含70个典型场景,揭示VLMs在移动操作中的空间定位挑战,强调增强空间感知的必要性。
中文摘要 AI 辅助
尽管视觉语言模型(VLMs)的进步赋予了机器人更强的环境理解和任务推理能力,但一个全面的评估方法对于推动VLMs在机器人导航和操作中的集成至关重要。然而,当前的基准缺乏评估多样化机器人任务的综合方法,且评估指标相对受限,这使得难以全面且细粒度地评估VLMs的具身能力。为解决这一问题,我们提出了RMMBench,一个评估基准,要求机器人理解语言指令并在连续空间中执行长时程任务。RMMBench将高层和低层具身任务无缝集成到一个统一框架中,构建了一个“导航-操作”任务套件,包含70个典型任务场景,范围从局部操作到长时程复合导航。结果表明,领先的VLMs在执行移动操作任务时仍面临空间定位的重大挑战,同时也凸显了在长时程交互中增强机器人空间感知能力的必要性。RMMBench可通过此https URL访问。
英文摘要
Although the advancement of vision-language models (VLMs) has endowed robots with enhanced environmental understanding and task reasoning, a comprehensive evaluation methodology is important to advance the integration of VLMs in robotic navigation and manipulation. However, current benchmarks lack a comprehensive method to evaluate diverse robotic tasks, and evaluation metrics remain relatively constrained, making it difficult to assess the embodied capabilities of VLMs in a thorough and fine-grained manner. To address this issue, we propose RMMBench, an evaluation benchmark that requires robots to understand language instructions and perform long-horizon tasks in continuous spaces. RMMBench seamlessly integrates high- and low-level embodied tasks into a unified framework, constructing a "navigation-manipulation" task suite comprising 70 canonical task scenarios that range from localized manipulation to long-horizon composite navigation. The results reveal that leading VLMs still face major challenges in spatial localization when performing mobile manipulation tasks, and also highlight the necessity of enhancing the spatial perception capability of robots during long-horizon interactions. RMMBench can be accessed at https://mxxq-stack.github.io/rmmbench-project/
发表机构
- Shandong University(山东大学)
- Meituan Group(美团集团)
机构由 AI 辅助整理,请以论文原文为准。