Verti-WM:面向越野强化学习的物理辅助外部感知世界模型
Verti-WM: A Physics-Aided Exteroceptive World Model for Off-Road Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对越野强化学习数据采集昂贵的问题,提出物理辅助外部感知世界模型Verti-WM,融合Transformer与地面力学模型,实现高效策略优化,预测误差降低34.6%,计算时间减少23.6倍,真实世界成功率80%。
中文摘要 AI 辅助
越野导航的强化学习需要大量的车辆-地形交互数据,而在高保真模拟器中收集这些数据成本高昂。世界模型提供了一种有前景的替代方案,通过在策略优化期间替代模拟器展开来降低数据需求。然而,越野世界模型必须基于外部感知的地形信息来调节状态转移,而仅靠本体感觉无法提供这些信息。此外,由于需要同时建模刚性地形和可变形地形,这一挑战进一步加剧,而数据驱动和基于物理的方法在此方面具有互补优势。我们提出了Verti-WM,一种物理辅助的外部感知世界模型,它通过循环融合一个用于刚性地形的冻结Transformer和一个用于可变形地形的神经符号地面力学模型。在每个预测位姿处,从提供的地图中查询高程和语义观测,以调节融合过程,从而无需进一步访问模拟器即可实现六自由度展开以进行策略优化。Verti-WM将预测误差分别比数据驱动基线和基于物理的基线降低了34.6%和21.7%。完全在Verti-WM内训练的策略达到了与直接在高保真模拟器中训练相当的任务成功率,同时计算时间减少了23.6倍。我们进一步使用真实世界数据验证了Verti-WM,使得在学习的真实世界运动动力学内进行策略优化成为可能,并在Verti-4-Wheeler平台上实现了80%的成功率,而直接进行模拟到现实迁移的成功率为40%。
英文摘要
Reinforcement learning for off-road navigation requires extensive vehicle-terrain interaction data, which are costly to collect in high-fidelity simulation. World models offer a promising alternative by replacing simulator roll-outs during policy optimization. However, an off-road world model must condition state transitions on exteroceptive terrain information, which proprioception alone does not provide. This challenge is further amplified by the need to model both rigid and deformable terrain, where data-driven and physics-based approaches offer complementary strengths. We propose Verti-WM, a physics-aided exteroceptive world model that recurrently fuses a frozen Transformer for rigid terrain and a neuro-symbolic terramechanics model for deformable terrain. Elevation and semantic observations queried from a supplied map at each predicted pose condition fusion, enabling six-degree-of-freedom rollouts for policy optimization without further simulator access. Verti-WM reduces prediction error by 34.6% and 21.7% over data-driven and physics-based baselines, respectively. Policies trained entirely within Verti-WM achieve comparable task success rates while reducing computation time by 23.6X relative to direct training in the high-fidelity simulator. We further validate Verti-WM using real-world data, enabling policy optimization within learned real-world kinodynamics and achieving a 80% success rate on the Verti-4-Wheeler platform, compared with 40% for direct sim-to-real transfer.
发表机构
- George Mason University(乔治梅森大学)
机构由 AI 辅助整理,请以论文原文为准。