CARO:适用于零样本鲁棒四足运动的接触无关残差观测
CARO: Contact-Agnostic Residual Observation for Zero-Shot Robust Quadruped Locomotion
- School of Aeronautic Science and Engineering, Beihang University(北京航空航天大学航空科学与工程学院)
- School of Automation Science and Electrical Engineering, Beihang University(北京航空航天大学自动化科学与电气工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出CARO框架,将固定基座欧拉-拉格朗日模型嵌入强化学习控制循环,无需额外传感器即可提升四足机器人零样本鲁棒性,在多项任务中表现出显著优势。
AI中文摘要:
我们提出了CARO,一种用于策略适配的接触无关残差观测框架。CARO将固定基座欧拉-拉格朗日模型嵌入强化学习控制循环,构建力矩级残差观测,无需力矩传感器、显式接触估计或基于视觉的浮动基座位置与线速度测量。干扰观测器提取表示动力学失配的结构化信号,策略学习利用该反馈进行在线适配。CARO在与标称策略相同的地形、指令和域随机化条件下训练,无需专门的干扰课程或额外适配监督,却在分布外有效载荷、质心偏移、地形几何、突发动力学变化及高架平台着陆等模拟和仿真到真实迁移任务中,实现了大幅提升的零样本鲁棒性。
英文摘要:
We propose CARO, a contact-agnostic residual observation framework for policy adaptation. CARO embeds a fixed-base Euler--Lagrange model into the reinforcement learning control loop and constructs a torque-level residual observation without requiring torque sensors, explicit contact estimation, or vision-based measurements of the floating-base position and linear velocity. A disturbance observer extracts a structured signal representing dynamics mismatch, while the policy learns to exploit this feedback for online adaptation. CARO is trained under the same terrain, command, and domain-randomization conditions as the nominal policy, without specialized disturbance curricula or additional adaptation supervision. Nevertheless, it achieves substantially improved zero-shot robustness in simulation and sim-to-real transfer tasks involving out-of-distribution payloads, center-of-mass shifts, terrain geometries, abrupt dynamics changes, and elevated-platform landings.