发表机构
Robotics and AI Institute (RAI); Robotics Institute, Carnegie Mellon University(机器人与人工智能研究所(RAI); 卡内基梅隆大学机器人研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于正则化记忆的在线学习框架,使双臂机器人在不到5分钟内安全掌握五种三球杂耍模式,核心是利用不完美先验知识实现高效稳定的真实世界动态操作技能学习。
AI 中文摘要
我们提出了一种在线学习框架,该框架即使在存在显著的仿真到现实(sim2real)差距的情况下,也能使双臂机器人在数分钟内直接在物理硬件上掌握多种杂耍模式。本研究最重要的经验之一是,即使模型与现实相差甚远,它也能对学习极为有用。这为我们方法的核心理念提供了动机:学习应建立在机器人当前的知识基础上,而非取代它。我们的基于正则化记忆的学习方法将这一原则付诸实践,通过从积累的经验中学习局部模型,同时保留全局先验模型,以便在经验稀疏的地方进行外推。这使得机器人能从每一次新经验中进行高效且稳定的在线学习,无需在庞大的可能行为空间中进行无信息探索。对于持续的机器人在线学习而言,安全性同样重要,它允许机器人在真实世界中反复练习和改进。我们构建了一个相互可达集,该集允许在连续的抛球和接球之间进行安全过渡,而不会使任一机械臂进入其后续动作需要违反机器人关节或执行器极限的状态。综合这些思路,配备多指手和机载视觉的双臂机器人能够在不到5分钟的真实世界交互时间内,安全地学习并组合五种典型的三球杂耍模式,包括 cascade( cascade杂耍)、tennis(网球杂耍)、half-shower(半淋浴杂耍)、shower(淋浴杂耍)和 box(盒式杂耍)。更广泛地说,本研究为机器人指明了方向:机器人可基于不完美的先验知识,通过自身的真实世界经验不断完善其行为。
英文摘要
We present an online learning framework that enables a bimanual robot to acquire diverse juggling patterns directly on physical hardware within minutes, even with a significant sim2real gap. One of the most important lessons from this work is that a model, even when far from reality, can be extremely useful for learning. This motivates a central philosophy of our approach: learning should build upon the robot's current knowledge rather than replace it. Our regularized memory-based learning puts this principle into practice by learning a local model from accumulated experience while retaining the global prior model to extrapolate where experience is sparse. This enables efficient and stable online learning from each new experience without resorting to uninformed exploration over a vast space of possible behaviors. Equally important to continual on-robot learning is safety, allowing the robot to repeatedly practice and improve in the real world. We construct a mutually reachable set that allows safe transitions between successive throws and catches, without driving either arm into a state from which its next action would require violating the robot's joint or actuator limits. Together, these ideas enable a bimanual robot with multi-fingered hands and onboard vision to safely learn and compose five canonical three-ball juggling patterns, including cascade, tennis, half-shower, shower, and box, within less than 5 minutes of real-world interaction. More broadly, this work points toward robots that build upon imperfect prior knowledge and continually refine their behavior through their own real-world experience.