发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlashNeRD通过并行流式架构、任意接触编码器和双向耦合,解决了NeRD的三大局限,实现高达55倍动力学加速、34秒训练ANYmal策略,并显著提升接触丰富任务性能。
AI 中文摘要
与解析物理相比,学习动力学模型有望实现更快、固有可微且易于适应真实数据的机器人仿真。神经机器人动力学(NeRD)通过保持碰撞检测的解析性,并用学习模型替代机器人的数值动力学来实现这一目标。但仍存在三个局限性:NeRD相对于其所学习的仿真器提速有限,仅在预定义点接受接触,并且与其操作的物体缺乏双向耦合。FlashNeRD通过并行流式架构消除了这三个问题,该架构使每次预测更快更准确,采用可接受任意位置接触的编码器,并与解析求解器仿真的物体实现双向耦合。在五个机器人上的实验表明,动力学更快更准确,策略学习更快,推理时规划也更快。在三个机器人上,FlashNeRD的动力学模型比优化后的NeRD快高达55倍,且在长时程滚动中更准确。这种加速延伸到策略学习,PPO在34秒内训练出ANYmal运动策略,比解析仿真器快3.8倍,比优化后的NeRD快2.7倍。使用基于采样的MPC方法DIAL-MPC,机器人攀爬了全部六个测试平台,而固定接触的NeRD仅能攀爬一个,规划时间约为解析仿真器的一半。使用FlashNeRD训练的立方体重定向策略完成度在仿真器训练策略目标数量的2%以内。
英文摘要
Compared with analytical physics, learned dynamics models promise robot simulation that is faster, inherently differentiable, and easily adaptable to real data. Neural Robot Dynamics (NeRD) pursues this by keeping collision detection analytical and replacing a simulator's numerical dynamics for the robot with a learned model. Three limitations remain. NeRD offers little speedup over the simulator it learned from, accepts contact only at predefined points, and has no two-way coupling with objects it manipulates. FlashNeRD removes all three with a parallel streaming architecture that makes each prediction faster and more accurate, an encoder that accepts contacts wherever they occur, and two-way coupling with objects simulated by analytical solvers. Experiments across five robots show faster and more accurate dynamics, faster policy learning, and faster inference-time planning. Across three robots, FlashNeRD's dynamics model is up to $55\times$ faster than an optimized NeRD and more accurate over long rollouts. This speedup extends to policy learning, where PPO trains an ANYmal locomotion policy in 34 seconds, $3.8\times$ faster than the analytical simulator and $2.7\times$ faster than an optimized NeRD. With DIAL-MPC, a sampling-based MPC method, the robot climbs all six test platforms where fixed-contact NeRD manages one, at approximately half the analytical simulator's planning time. Cube-reorientation policies trained with FlashNeRD complete within 2% of the simulator-trained policy's target count.