arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

速度中的精度:用于液压挖掘机控制的样本高效在线基于模型的强化学习

Precision at Speed: Sample-Efficient Online Model-Based Reinforcement Learning for Hydraulic Excavator Control

Claudio Canales, Fang Nan, Marco Hutter, Javier Ruiz-del-Solar

arXiv 2609.31025首次发表:更新:

发表机构

ETH Zürich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种在线基于模型的强化学习框架,用于液压挖掘机控制,通过概率动力学集成和精度门控目标,在无演示或预训练下实现样本高效学习,20分钟达到先前100-150分钟训练数据的精度,40分钟维持亚厘米级误差。

AI 中文摘要

对于具有复杂驱动动力学的机器人,精确、高速的控制仍然具有挑战性。直接在硬件上学习进一步受到现实世界交互成本的限制。我们提出了一种在线基于模型的强化学习框架,该框架从头学习一个概率动力学集成模型,用于基于采样的模型预测控制。一个精度门控的轮廓目标将进度奖励条件化于路径精度,优先考虑精度而非速度。在数据驱动的挖掘机模拟器中,该框架比评估的基于模型的强化学习基线实现了更高的样本效率。我们通过在11.5吨Menzi Muck M445液压挖掘机上直接学习来验证该框架,无需演示或模拟预训练。经过20分钟的交互,控制器达到了与先前在100-150分钟数据上训练的 learned 控制器相当的跟踪精度。经过40分钟,它能在高运行速度下维持亚厘米级的平均路径误差。

英文摘要

Precise, high-speed control remains challenging for robots with complex actuation dynamics. Learning directly on hardware is further constrained by the cost of real-world interaction. We present an online model-based reinforcement learning framework that learns a probabilistic dynamics ensemble model from scratch for sampling-based model predictive control. A precision-gated contouring objective conditions the progress reward on path accuracy, prioritizing precision over speed. In a data-driven excavator simulator, the framework achieves higher sample efficiency than the evaluated model-based reinforcement learning baselines. We validate the framework by learning directly on an 11.5-ton Menzi Muck M445 hydraulic excavator, without demonstrations or simulation pretraining. After 20 minutes of interaction, the controller reaches tracking accuracy comparable to prior learned controllers trained on 100-150 minutes of data. After 40 minutes, it sustains sub-centimeter mean path error at high operating speeds.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑