基于最少数据的通用机器人策略适配
Adaptation of Generalist Robot Policies with Minimal Data
浏览论文内容
中文总结 AI 辅助
本文提出MiDAS算法,结合离线行为克隆与在线强化学习,实现机器人策略仅用1次演示即可完成任务适配,性能优于基线且能泛化,为机器人自主学习提供可行方案。
中文摘要 AI 辅助
机器人学习的核心目标是超越特定任务的人工数据收集,实现机器人通过自主交互提升性能。但当前策略下完全自主学习仍具难度:稀疏奖励与零样本探索能力薄弱,导致机器人难以从零发现成功行为。本文研究最少数据适配,即预训练机器人策略需从仅1次演示及后续自主在线交互中学习新任务,该设置是完全自主改进的最接近可行代理,可探究最少人工引导能否启动自主学习及所需算法要素。本文构建MiDAS,一种简单的离线到在线强化学习方案:首先通过单/少量演示的行为克隆将预训练VLA锚定到目标任务,再通过残差策略参数化的基于值的在线强化学习优化。在LIBERO和RoboCasa数据集上,MiDAS仅用1次演示即可恢复强任务性能,大幅优于基线方法且能泛化到演示外条件。本文还在双机械臂YAM平台评估MiDAS:从1次演示得到的低成功率脆弱策略出发,MiDAS在约6小时在线交互中提升策略鲁棒性并学习到新的成功行为。据所知,这是首次实现从单任务演示可靠适配机器人策略的验证。
英文摘要
A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction. Yet fully autonomous learning remains difficult with current policies: sparse rewards and weak zero-shot exploration make it unlikely that a robot will discover successful behavior from scratch. We study minimal-data adaptation, a regime in which a pre-trained robot policy must learn a new task from as little as one demonstration followed by autonomous online interaction. This setting serves as the closest tractable proxy for fully autonomous improvement, allowing us to study whether minimal human guidance can bootstrap autonomous learning and what algorithmic ingredients make it feasible. We build MiDAS, a simple offline-to-online RL recipe that first anchors a pre-trained VLA to the target task with behavior cloning on single/few demonstrations, then improves it through value-based online RL on a residual policy parameterization. Across LIBERO and RoboCasa, MiDAS recovers strong task performance from as little as one demonstration, substantially outperforming baselines and generalizing beyond demonstrated conditions. We further evaluate MiDAS on a bimanual YAM platform. Starting from a fragile low-success policy obtained from a single demonstration, MiDAS improves its robustness and learns new successful behaviors over ~6 hours of online interaction. To the best of our knowledge, this is the first demonstration of reliable robot policy adaptation from a single task demonstration.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。