发表机构
LAAS-CNRS, Université de Toulouse, CNRS; Machines in Motion Laboratory, New York University; Industrial Engineering Department, University of Trento; Artificial and Natural Intelligence Toulouse Institute (ANITI)(LAAS-CNRS,图卢兹大学,法国国家科学研究中心; 纽约大学运动机器实验室; 特伦托大学工业工程系; 图卢兹人工与自然智能研究所(ANITI))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VAMPS利用MPPI控制,无需人工演示即可在仿真或真实数据上训练可复用机器人策略,支持运动与视觉运动任务,并成功迁移到硬件。
AI 中文摘要
直接在物理系统上学习机器人策略仍然困难,因为数据收集成本高昂且策略探索可能不安全。我们提出了基于采样规划的视觉与运动策略(VAMPS),这是一个利用模型预测路径积分(MPPI)控制来训练可复用策略的框架,无需人工演示。VAMPS支持两种训练模式。对于单步本体感觉策略,它在仿真中迭代运行:策略为MPPI提供热启动,随着策略变化,优化后的轨迹提供新的监督。学习到的终端价值改善了短视界规划,而隐式Q学习(IQL)评论器指导策略更新。迭代细化优于在冻结的MPPI数据上仅训练一次,我们将学习到的运动策略迁移到宇树Go2。对于视觉运动策略,VAMPS直接从真实机器人数据中运行。MPPI使用任务特定的状态估计来规划和执行轨迹,同时在Flexiv Rizon 10S上记录RGB和传感器观测。带Transformer的动作分块预测动作块,减少了有效预测视界,并在该固定数据集上离线训练。我们展示了视觉运动抓取放置和力感知白板擦除。在后一项任务中,策略额外观测测量的$6$维力和期望法向力。这些结果表明,VAMPS可以在仿真中学习策略后迁移到硬件,或直接从自主收集的真实机器人数据中学习策略。
英文摘要
Learning robot policies directly on physical systems remains difficult because data collection is costly and policy exploration can be unsafe. We introduce Visual and Motor Policies from Sampling-Based Planning (VAMPS), a framework that uses Model Predictive Path Integral (MPPI) control to train reusable policies without human demonstrations. VAMPS supports two training modes. For one-step proprioceptive policies, it operates iteratively in simulation: the policy warm-starts MPPI, and the refined trajectories provide new supervision as the policy changes. A learned terminal value improves short-horizon planning, while an Implicit Q-Learning (IQL) critic guides the policy update. Iterative refinement outperforms training once on frozen MPPI data, and we transfer the learned locomotion policy to a Unitree Go2. For visuomotor policies, VAMPS operates directly from real-robot data. MPPI uses task-specific state estimates to plan and execute trajectories while recording RGB and sensor observations on a Flexiv Rizon 10S. Action Chunking with Transformers predicts action chunks, reducing the effective prediction horizon, and is trained offline on this fixed dataset. We demonstrate visuomotor pick-and-place and force-aware whiteboard erasing. In the latter task, the policy additionally observes the measured $6$-D wrench and desired normal force. These results show that VAMPS can learn policies either in simulation followed by hardware transfer or directly from autonomously collected real-robot data.