发表机构
ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种在线基于模型的强化学习框架,用于液压挖掘机控制,通过概率动力学集成和精度门控目标,在无演示或预训练下实现样本高效学习,20分钟达到先前100-150分钟训练数据的精度,40分钟维持亚厘米级误差。
AI 中文摘要
对于具有复杂驱动动力学的机器人,精确、高速的控制仍然具有挑战性。直接在硬件上学习进一步受到现实世界交互成本的限制。我们提出了一种在线基于模型的强化学习框架,该框架从头学习一个概率动力学集成模型,用于基于采样的模型预测控制。一个精度门控的轮廓目标将进度奖励条件化于路径精度,优先考虑精度而非速度。在数据驱动的挖掘机模拟器中,该框架比评估的基于模型的强化学习基线实现了更高的样本效率。我们通过在11.5吨Menzi Muck M445液压挖掘机上直接学习来验证该框架,无需演示或模拟预训练。经过20分钟的交互,控制器达到了与先前在100-150分钟数据上训练的 learned 控制器相当的跟踪精度。经过40分钟,它能在高运行速度下维持亚厘米级的平均路径误差。
英文摘要
Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learns a probabilistic dynamics ensemble model from scratch for sampling-based model predictive control. A precision-gated contouring objective conditions the progress reward on path accuracy, prioritizing precision over speed. In a data-driven excavator simulator, the framework achieves higher sample efficiency than the evaluated model-based reinforcement learning baselines. We validate the framework by learning directly on an 11.5-ton Menzi Muck M445 hydraulic excavator, without demonstrations or simulation pretraining. After 20 minutes of interaction, the controller reaches tracking accuracy comparable to prior learned controllers trained on 100-150 minutes of data. After 40 minutes, it sustains sub-centimeter mean path error at high operating speeds.