arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18732cs.RO

PASSAGE:面向杂乱环境中具身感知人形机器人穿越的场景对齐运动学习规模化

PASSAGE: Scaling Scene-Aligned Motion Learning for Perceptive Humanoid Traversal in Cluttered Environments

  • Galbot
  • Shanghai Qi Zhi Institute(上海期智研究院)
  • ShanghaiTech University(上海科技大学)
  • Zhongguancun Academy(中关村学院)
  • Shanghai Jiao Tong University(上海交通大学)
  • National University of Singapore(新加坡国立大学)
  • Tsinghua University(清华大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuxuan Ma, Zicheng Zeng, Chunlin Peng, Zhoujian Li, Zetong Zhao, Zhikai Zhang, Yunrui Lian, Han Xue, Sikai Liang, Weiyi Zhu, Mulin Chen, Chenghuai Lin, Jiayu Ze… 展开作者

Yuxuan Ma, Zicheng Zeng, Chunlin Peng, Zhoujian Li, Zetong Zhao, Zhikai Zhang, Yunrui Lian, Han Xue, Sikai Liang, Weiyi Zhu, Mulin Chen, Chenghuai Lin, Jiayu Zeng, Yanwei An, Songan Zhang, Jiayuan Gu, Jilong Wang, Jingbo Wang, He Wang, Li Yi

AI总结:

PASSAGE提出感知条件规划器-跟踪器框架,利用100小时场景对齐运动数据,通过条件流匹配与全身跟踪实现人形机器人在杂乱环境中的自主穿越,成功率提升至70.3%。

AI中文摘要:

人形机器人可以跨过、挤过和低头钻过障碍物,但从机载感知中学习选择并协调这些行为仍然具有挑战性。许多现有方法依赖于特定任务强化学习目标或精心策划的运动库,导致广泛的行为覆盖成本高昂。我们提出PASSAGE,一个用于人形机器人穿越的感知条件规划器-跟踪器框架。利用虚拟现实和惯性动作捕捉,我们在1,500个杂乱场景中收集了100小时的场景对齐人体运动。条件流匹配规划器从运动历史、局部目的地和以机器人为中心的多层高程图生成短时域参考,而感知全身跟踪器以50Hz频率执行这些参考并带有几何反馈。实时分块促进了块间一致性,在冻结跟踪器下的规划器侧强化学习后训练进一步改善了闭环性能。无需技能标注或障碍物特定策略,一个规划器-跟踪器对即可在未见几何形状中选择并组合穿越行为。在仿真中,组件消融量化了每个阶段的贡献。在三个独立训练种子下,将捕获数据从6小时扩展到100小时,在保留场景上平均无接触成功率从48.1%提升至68.9%,而经过验证场景增强的最终模型达到70.3%。全机载系统在Jetson AGX Orin上集成了以自我为中心的3D LiDAR感知、在线占用地图构建、6.25Hz规划和50Hz控制;在50个未见物理布局上的测试表明,无需预建地图或机外计算即可实现穿越。

英文摘要:

Humanoid robots can step over, squeeze past, and duck under obstacles, but learning to select and coordinate these behaviors from onboard perception remains challenging. Many existing approaches rely on task-specific reinforcement-learning objectives or curated motion libraries, making broad behavioral coverage costly. We present PASSAGE, a perception-conditioned planner--tracker framework for humanoid traversal. Using virtual reality and inertial motion capture, we collect 100 h of scene-aligned human motion across 1,500 cluttered scenes. A conditional flow-matching planner generates short-horizon references from motion history, a local destination, and a robot-centric multi-layer elevation map, while a perceptive whole-body tracker executes them at 50 Hz with geometric feedback. Real-time chunking promotes inter-chunk consistency, and planner-side RL post-training under the frozen tracker further improves closed-loop performance. Without skill annotations or obstacle-specific policies, one planner--tracker pair selects and composes traversal behaviors across unseen geometries. In simulation, component ablations quantify the contribution of each stage. Across three independent training seeds, scaling captured data from 6 to 100 h increases mean contact-free success from 48.1% to 68.9% on held-out scenes, while the final model with validated scene augmentation reaches 70.3%. The fully onboard system integrates egocentric 3D LiDAR perception, online occupancy mapping, 6.25 Hz planning, and 50 Hz control on a Jetson AGX Orin; tests across 50 unseen physical layouts demonstrate traversal without prebuilt maps or offboard computation.

相关深度报道

↑