发表机构
Stanford University; Harvard University; Stanford University School of Medicine(斯坦福大学; 哈佛大学; 斯坦福大学医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对低成本四足硬件的仿真-现实差距,采用生物启发方法,结合延迟前向模型与时间感知神经网络,学习到鲁棒CPG,实现了成本受限硬件上的强化学习控制。
AI 中文摘要
将学习到的控制策略部署到低成本机器人平台时,会引入传输延迟和噪声电机反馈,系统性地扩大了仿真到现实的差距。硬件中仿真到部署的鸿沟在于执行器到达指令位置的延迟,在Mini Pupper 2这类平台上,实测超过50毫秒的传输延迟会将运动任务从标准马尔可夫决策过程转变为部分可观测的过程。本文采用受生物启发的方法处理噪声和延迟反馈以缩小仿真到现实的差距,从而扩展强化学习在成本受限硬件上的能力。使用低成本四足硬件平台,我们发现结合平均执行器延迟前向模型与时间感知神经网络可实现鲁棒运动。此外,该时间感知神经网络学习到了中枢模式发生器(CPG),即一种能对+320毫秒延迟扰动具有鲁棒性的自维持节律步态,这与脊椎动物脊髓中的CPG类似。我们认为时间自组织可能是成本受限运动的通用策略。
英文摘要
Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. The chasm of simulation to deployment in hardware lies in the delay of the actuator reaching the commanded position. On platforms such as the Mini Pupper 2, a measured >50 ms transport delay transforms the locomotion task from a standard Markov decision process into a partially observable one. In this paper, we take a biologically inspired approach of handling noisy and delayed feedback to close the sim-to-real gap, thereby expanding the capability of reinforcement learning on cost-constrained hardware. Using a low-cost quadrupedal hardware platform, we find that using a forward model of the average actuator delay, paired with a time-aware neural network results in robust locomotion. Additionally, our time-aware neural network learned a central pattern generator (CPG): a self-sustaining rhythmic gait that is robust to +320 ms latency perturbations, mirroring the CPGs found in the spinal cords of vertebrates. We posit that temporal self-organization may be a general strategy for cost-constrained locomotion.
CommentsSim-to-real transfer, locomotion, reinforcement learning, central pattern generator