arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习正确的抽象:用于复杂机器人控制的神经简化动力学

Learning the Right Abstraction: Neural Reduced Dynamics for Complex Robot Control

Harry Zhang, Dan Negrut

arXiv 2608.19375首次发表:更新:

发表机构

University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出神经简化动力学(NRD)框架,学习保留控制相关物理特性的简化状态,训练的策略可迁移至高保真模拟器,在三项控制任务中实现高精度且速度大幅提升的机器人控制。

AI 中文摘要

高保真具身AI模拟器能为复杂机器人系统提供逼真的评估,但其计算成本限制了其直接用于大规模强化学习任务。本文提倡使用精度稍低但速度更快的模拟方法,这类方法可采用数据驱动模型,例如神经动力学模型。本文指出,神经动力学模型在复杂机器人控制中的实用价值在于学习“正确的抽象”:一种简化状态,该状态既保留了高保真系统中与控制相关的物理特性,又能实现高吞吐量的策略学习。我们开发了神经简化动力学(NRD)框架,该框架将模型传播的状态与可作为输入提供或通过解析恢复的状态分离开,在冻结的学习模型内部完整训练策略,并在高保真模拟器中对这些策略进行验证。通过两个案例研究,我们在三项控制任务中实例化了该框架:在刚性、崎岖和可变形的连续体表示模型(CRM)地形上的地形感知HMMWV轨迹跟踪;以及普通履带式车辆及其前端安装的铰接式机械臂的目标到达。所有策略均可迁移回高保真模拟器。在地形条件动力学模型内部训练的单个策略,在没有自身地形输入的情况下,在所有三种地形上的中值和平均跟踪误差均低于单地形专家策略,包括零样本崎岖地形。量化结果显示,履带式车辆在100个目标中成功到达100个,机械臂在100个目标中成功到达97个,且无任何接触或关节极限违规情况。NRD模型的模拟时间速度比其替代的高保真模拟器场景快约四个数量级,使迭代式在线策略学习成为可能,并支持神经简化动力学作为精确但昂贵的物理模拟与可扩展机器人学习之间的桥梁。

英文摘要

High-fidelity embodied AI simulators provide realistic evaluation of complex robotic systems, but their computational cost limits their direct use for large-scale reinforcement learning campaigns. We advocate the use of less accurate but more expeditious simulations, which might draw on data-driven, e.g., neural dynamics, models. This contribution argues that the practical value of a neural dynamics model for complex robot control lies in learning the \emph{right abstraction}: a reduced state that preserves the control-relevant physics of the high-fidelity system while enabling high-throughput policy learning. We develop a neural reduced dynamics (NRD) framework that separates the state the model propagates from what can be supplied as an input or recovered analytically, trains policies entirely inside the frozen learned model, and validates them back in the high-fidelity simulator. Two case studies instantiate it across three control tasks: terrain-aware HMMWV trajectory tracking on rigid, bumpy and deformable Continuum Representation Model (CRM) terrain; and goal reaching for a stock tracked vehicle and its front-mounted articulated arm. Every policy transfers back to the high-fidelity simulator. A single policy trained inside the terrain-conditioned dynamics model, and given no terrain input of its own, attains lower median and mean tracking error than both single-terrain specialists on all three terrains, including zero-shot bumpy terrain. Quantitatively, the tracked vehicle reaches 100 of 100 goals and the arm 97 of 100, with zero contacts or joint-limit violations. The NRD models advance roughly four orders of magnitude faster in simulated time than the high-fidelity simulator scenes they replace, making iterative on-policy learning practical and supporting neural reduced dynamics as a bridge between accurate but expensive physics simulation and scalable robot learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑