X-Reset:通过跨具身重置扩展以物体为中心的强化学习
X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets
- Applied Intuition
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
X-Reset框架通过将人类手-物体演示运动学重定向为噪声机器人状态并采样为重置,解决以物体为中心的强化学习探索难题,在三个具身上对20个物体训练通用策略,并实现零样本迁移。
中文摘要 AI 辅助
仿真中的强化学习(RL)可以在没有机器人演示的情况下训练灵巧操作策略,但使用与任务无关的奖励训练单一通用策略面临严重的探索问题:接近、抓取和重新定向具有多个自由度的多样物体难以从零开始发现。先前的工作通过高质量的机器人演示、每任务奖励塑形或限制策略为狭窄行为模式来使探索可行。我们提出X-Reset,一个通过人类手-物体演示解决探索问题的框架。X-Reset不是模仿或跟踪重定向的人类运动,而是将手-物体状态运动学重定向到带噪声的机器人状态,过滤掉在仿真中不稳定的状态,并在使用通用物体中心奖励的RL训练期间将剩余状态采样为重置。所得策略仅依赖于物体状态和目标,演示通过重置分布进入训练。我们展示了X-Reset在三个具身(一个22自由度手连接在两个不同手臂和一个平行爪夹持器)上对20个物体训练通用策略,并解决了从零开始RL的探索挑战。X-Reset随训练物体数量扩展,泛化到未见物体,可以从不完美的手姿态估计中学习,并将行为从仿真到现实零样本迁移。
英文摘要
Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.