发表机构
The University of Manchester; X-Humanoid; University of Warwick; Newcastle University(曼彻斯特大学; X-Humanoid; 华威大学; 纽卡斯尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对杂乱室内环境中人形机器人长时程全身操作-移动任务,提出Humanoid Horizon框架,通过并行训练、动态起点和奖励门控机制,在双物体基准上各阶段成功率超80%,并缓解了长地平线下的性能退化。
AI 中文摘要
杂乱的室内环境中,大型和重型物体散布在不同表面上,要求人形机器人在一个不间断的连续情节中依次导航、抓取、运输并准确放置每个物品到目标位置。这种长时程、全身操作-移动任务对当前方法来说仍然是一个重大挑战。以往的方法常常面临两个主要问题:易奖励偏差,即训练过度强调早期运输阶段而忽视后期阶段;以及灾难性遗忘,即专注于后期阶段导致早期阶段性能下降。在本工作中,我们提出了人形地平线(Humanoid Horizon),一个统一策略框架,通过三个相互关联的机制来克服这些限制。并行训练策略将N个场景组织成S个并发的阶段流,由共享策略控制,确保所有运输阶段获得连续的梯度更新,消除了顺序优化的瓶颈。动态起点机制利用上游rollout的终端状态更新每个环境的初始状态,逐步扩大过渡覆盖范围,增强阶段边界的鲁棒性。奖励门控机制在后期阶段流中,当紧邻的前一个物体被移动超过设定阈值时,将本情节剩余部分的奖励设为零,从而让共享策略学会不干扰刚放置的物体,并确保整个情节中早期放置得以保留。综合这些策略,在双物体LHM-Humanoid基准(350个训练场景,66个保留场景)上,各阶段成功率超过80%。随着顺序运输物体数量超过两个,成功率随地平线增加而下降,但相对于所有基线的急剧下降,这种退化是平缓的。
英文摘要
Cluttered indoor environments, where large and heavy objects are scattered across diverse surfaces, require humanoid robots to sequentially navigate, grasp, transport, and accurately place each item at its target location within a single uninterrupted episode. This long-horizon, whole-body loco-manipulation task remains a significant challenge for current methods. Previous approaches often suffer from two main issues: easy-reward bias, where training overemphasizes early transport stages at the expense of later ones, and catastrophic forgetting, where focusing on later stages leads to a decline in earlier-stage performance. In this work, we introduce Humanoid Horizon, a unified policy framework designed to overcome these limitations through three interrelated mechanisms. The Parallel Training Strategy organizes $N$ scenes into $S$ concurrent stage streams governed by a shared policy, ensuring all transport stages receive continuous gradient updates and removing the bottleneck of sequential optimization. The Dynamic Starting Mechanism updates each environment's initial state with terminal states from upstream rollouts, gradually broadening transition coverage and enhancing robustness at stage boundaries. Reward Gating sets the reward to zero for the rest of the episode in later-stage streams when the immediately preceding object is displaced beyond a set threshold, so the shared policy learns not to disturb a just-placed object and earlier placements are preserved throughout the episode. Collectively, these strategies achieve per-stage success rates exceeding 80\% on the two-object LHM-Humanoid benchmark (350 training scenes, 66 held-out scenes). As the number of sequentially transported objects grows beyond two, success declines with the horizon, but the degradation is graceful relative to the sharp drop seen in all baselines.
CommentsVideos and results are available at https://haozhuo-zhang.github.io/Humanoid-Horizon-project-page/