发表机构
NVIDIA Corporation(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
WeaveRL提出GPU加速的surfel场景重建方法,集成到几何织物中,使强化学习策略能处理复杂几何操作,并在新障碍下将无碰撞完成率从35%提升至61%。
AI 中文摘要
强化学习使机器人能够获取复杂技能,但为几何复杂的操作生成策略仍然困难。一种有前景的方法是在避碰控制器(如几何织物)之上进行学习。然而,这些方法依赖于静态、手工指定的场景表示。将主动的在线3D感知集成到大规模并行强化学习训练中,迄今为止仍不可及。我们提出一种GPU加速方法,在主动 rollout 期间,跨数千个并行仿真实例将场景重建为 surfel 集合。这使得策略能够基于传感器派生的几何而非手工指定的几何进行操作。在一系列碰撞密集的操作任务中,我们的 surfel 织物使策略能够应对基于基元的基线方法失败的几何复杂场景,同时保持仿真到现实的迁移。此外,使用场景感知织物学习的策略在测试时对引入新几何更为鲁棒,将未见障碍物下的无碰撞任务完成率从35%提升至61%。我们发布重建系统、训练代码和测试数据集,以推动该方向的研究。
英文摘要
Reinforcement learning allows robots to acquire complex skills, but producing policies for geometrically complex manipulation remains difficult. A promising approach is to learn on top of collision-avoidant controllers, such as geometric fabrics. However, these approaches have relied on static, hand-specified representations of the scene. Integrating active, online 3D perception into massively parallel RL training has so far been inaccessible. We introduce a GPU-accelerated method that reconstructs the scene as a collection of surfels across thousands of parallel simulation instances during active rollouts. This lets policies operate over sensor-derived, rather than hand-specified, geometry. On a suite of collision-dense manipulation tasks, our surfel fabrics enable policies to tackle geometrically complex scenes where primitive-based baselines fail, while maintaining sim-to-real transfer. Furthermore, policies learned with a scene-aware fabric are more robust to the introduction of novel geometry at test time, improving collision-free task completion under unseen obstacles from 35% to 61%. We release our reconstruction system, training code and test dataset to spur research in this direction.