arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SOLO:稳定全地形长视距感知类人机器人运动

SOLO: Stable Omni-terrain Long-Horizon Perceptive Humanoid Locomotion

Pihai Sun, Gang Han, Jingkai Sun, Jiahao Ma, Zeran Su, Zelin Tao, Peiran Liu, Shuai Shi, Wei Cui, Zifan Wang, Jialin Yu, Wen Zhao, Kangning Yin, Jiaxu Wang, Jiahang Cao, Lingfeng Zhang, Hao Cheng, Jian Tang, Qiang Zhang, Yijie Guo

arXiv 2608.26583首次发表:更新:

发表机构

Artificial General Intelligence Institute, University of Science and Technology of China; X-Humanoid; The University of Hong Kong; The Australian National University; The Hong Kong University of Science and Technology (Guangzhou); Shanghai Jiao Tong University; The Chinese University of Hong Kong; Tsinghua University(中国科学技术大学通用人工智能研究院; X-人形机器人; 香港大学; 澳大利亚国立大学; 香港科技大学(广州); 上海交通大学; 香港中文大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SOLO是解决长视距类人机器人运动脆弱性的统一框架,通过QR与TA-MSE蒸馏提升地形感知与策略学习,仿真及实际测试中通行成功率显著优于基线方法,可零样本完成长距离复杂地形任务。

AI 中文摘要

人类可在长距离复杂地形中保持平衡行进,而感知类人机器人策略会因感知与控制误差累积变得脆弱。本文提出SOLO这一统一框架,解决导致长视距脆弱性的两个复合原因:密集地形重建平滑了对动作关键的细节,而逐点模仿缺乏时间信用分配。其查询重构器(QR)采用傅里叶编码的单元查询,从深度-本体感受令牌中检索空间特定证据,保留清晰地形边界;轨迹感知均方误差(TA-MSE)蒸馏将下一状态的师生分歧添加到PPO奖励中,使广义优势估计能将未来分歧惩罚传播到先前动作。在仿真中,QR将高度图L1误差降低3.3-4.0倍,TA-MSE在课程进度上优于PPO和MSE+PPO;在压力测试地形上,SOLO实现97.5%的平均通行成功率和96%的踏脚石成功率,而密集重构器变体分别为75.0-75.6%和0-3%。仅使用胸部安装的深度相机和本体感受零样本部署时,SOLO完成了1.5公里的连续室外路线和室内混合地形路线。项目页面:this https URL

英文摘要

Humans traverse complex terrain over long distances without losing balance, whereas perceptive humanoid policies become fragile as perception and control errors accumulate. We present SOLO, a unified framework addressing two compounding causes of this long-horizon fragility: dense terrain reconstruction smooths action-critical details, and pointwise imitation lacks temporal credit assignment. Its Query Reconstructor (QR) uses Fourier-encoded cell queries to retrieve spatially specific evidence from depth-proprioception tokens, preserving sharp terrain boundaries. Trajectory-Aware MSE (TA-MSE) Distillation adds next-state teacher-student disagreement to the PPO reward, enabling Generalized Advantage Estimation to propagate future disagreement penalties to preceding actions. In simulation, QR reduces height-map L1 error by factors of 3.3-4.0, while TA-MSE surpasses PPO and MSE+PPO in curriculum progression. On stress-test terrains, SOLO achieves 97.5% mean traversal success and 96% stepping-stone success, versus 75.0-75.6% and 0-3% for dense-reconstructor variants. Deployed zero-shot with only a chest-mounted depth camera and proprioception, SOLO completes a continuous 1.5-km outdoor route and an indoor mixed-terrain course. Project page: https://sunpihai-up.github.io/solo/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑