智能体场景策略:统一空间、语义与可供性用于机器人动作
Agentic Scene Policies
- Université de Montréal(蒙特利尔大学)
- Mila - Quebec AI Institute(魁北克人工智能研究院)
- Sapienza University of Rome(罗马萨皮恩扎大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出智能体场景策略(ASP),利用场景表示的语义、空间与可供性查询能力,实现零样本语言条件机器人策略,在桌面操作和房间级任务中优于视觉-语言-动作模型。
AI中文摘要:
执行开放式自然语言查询是机器人学的核心问题。尽管模仿学习和视觉-语言-动作模型(VLAs)的最新进展已经实现了有前景的端到端策略,但这些模型在面对复杂指令和新场景时仍会遇到困难。另一种方法是设计一个显式的场景表示,作为机器人与世界之间的可查询接口,利用查询结果来指导下游的运动规划。在这项工作中,我们提出了智能体场景策略(ASP),这是一个智能体框架,利用现代场景表示的先进语义、空间和基于可供性的查询能力来实现一个强大的语言条件机器人策略。ASP可以通过显式推理对象可供性来以零样本方式执行开放词汇查询,以应对更复杂的技能。通过大量实验,我们将ASP与VLAs在桌面操作问题上进行比较,并展示了ASP如何通过可供性引导的导航和扩展的场景表示来处理房间级查询。(项目页面:https://montrealrobotics.ca/agentic-scene-policies.github.io/)
英文摘要:
Designing or learning robot policies that generalize zero-shot across a range of language instructions and objects is a core problem in robotics. Vision-Language-Action models (VLAs) learn such policies end-to-end by repurposing existing Vision-Language Models (VLMs), but generalization to new instructions and objects remains challenging. An alternative is to implement a modular policy by leveraging an explicit VLM-based 3D scene representation and motion planning. While modular policies show strong zero-shot potential, they typically retrieve objects based on semantics without explicit spatial reasoning, severely restricting their overall grounding capabilities. They also interact with objects using basic grasping and navigation skills. In this work, we address these limitations by unifying grounding capabilities and robot skills in a single agentic action space through a scene-agent tool interface. By leveraging part-level affordances, our skills generalize across diverse objects and enable zero-shot interactions such as unplugging chargers and opening drawers. We name the resulting framework Agentic Scene Policies (ASP). Through extensive real-world experiments, we show how ASP consistently outperforms leading VLAs in the zero-shot setting. We also demonstrate the extensibility of our framework by introducing a mobile version of ASP to tackle room-level queries. See our project page (https://montrealrobotics.ca/agentic-scene-policies.github.io/) for more results.