arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39755cs.RO

经验驱动的四足机器人地形可通行性持续学习

Experience-Driven Continual Learning of Terrain Traversability for Quadruped Robots

  • Istituto Italiano di Tecnologia(意大利理工学院)
  • Università degli Studi di Genova(热那亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Luca Bricarello, João Carlos Virgolino Soares, Alberto Sanchez-Delgado, Fulvio Mastrogiovanni, Claudio Semini

AI总结:

本文提出一种基于运动经验的持续学习流程,利用冻结的DINOv3描述符预测四足机器人地形交互指标,并通过门控回放机制在适应新地形时减少旧知识退化,提升导航安全性与效率。

AI中文摘要:

在不熟悉的地形上进行安全高效的四足导航需要在接触前预测地形与机器人之间的相互作用:仅凭几何和视觉外观无法揭示机器人将如何打滑、加载其足部或消耗能量。本文提出了一种持续学习流程,利用运动经验从接触前图像中学习这些相互作用结果,并在观察到新接触时持续更新预测。由训练期间冻结的DINOv3骨干模型生成的接触前描述符被映射到五个根据测量可靠性加权的本体感觉指标:平面足部打滑、平均法向地面反作用力、牵引指数、运输成本以及触地加载速率。一个紧凑的证据回归器使我们能够从视觉描述符中预测这些指标以及偶然不确定性和认知不确定性。持续适应结合了有界经验回放与验证门:候选模型仅在改进最近保留数据上的性能同时将历史保留数据上的退化保持在规定容差内时,才替换部署的预测器。预测和认知不确定性被投影到局部多层地图中,并组合成保守的可通行性得分地图,其属性权重可以在不重新训练的情况下调整。生成的地图用于下游导航测试。ROS2实现支持在仿真和硬件上对Unitree Go2进行评估,模型在每个领域分别训练。在三个先前未见地形的顺序硬件流上,门控回放相对于无门回放将最终锚点负对数似然(NLL)退化减少了23.1%,同时实现了相似的新地形适应。

英文摘要:

Safe and efficient quadruped navigation over unfamiliar terrain requires predicting terrain-robot interaction before contact: geometry and visual appearance alone cannot reveal how the robot will slip, load its feet, or expend energy. This paper presents a continual learning pipeline that uses locomotion experience to learn these interaction outcomes from pre-contact images and continually updates the predictions as new contacts are observed. Pre-contact descriptors, produced by a DINOv3 backbone model frozen during training, are mapped to five proprioceptive indicators weighted according to measurement reliability: planar foot slip, mean normal ground-reaction force, traction index, cost of transport, and touchdown loading rate. A compact evidential regressor allows us to predict these indicators together with aleatoric and epistemic uncertainty from the visual descriptors. Continual adaptation combines bounded experience replay with a validation gate: candidate models replace the deployed predictor only when they improve performance on recent held-out data while keeping degradation on historical held-out data within a prescribed tolerance. Predictions and epistemic uncertainty are projected into a local multilayer map and combined into a conservative traversability score map whose property weights can be adjusted without retraining. The resulting map is used for downstream navigation tests. The ROS2 implementation supports evaluation on a Unitree Go2 in simulation and on hardware, with models trained separately in each domain. On a sequential hardware stream over three previously unseen terrains, gated replay reduces final anchor negative log-likelihood (NLL) degradation by 23.1% relative to replay without the gate while attaining similar new-terrain adaptation.

↑