arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05999cs.RO

超越扁平策略:面向机器人操作具身智能体的分层后训练

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation

He Kong, Zengjue Chen, Qi Wang, Qianli Xing, Runliang Niu, Peidong Liu, Jiawei Li, Shiqi Wang, Yi Chang

AI总结:

针对现有VLA模型扁平策略难以处理长时程操作的问题,提出HiRoC分层后训练框架,通过解耦规划与执行并对齐分布,在机器人操作基准上取得优于强基线的性能。

AI中文摘要:

视觉-语言-动作(VLA)模型通过利用预训练视觉-语言模型,在机器人操作领域展现出卓越能力。然而,现有后训练方法大多将VLA模型优化为扁平策略,难以显式建模任务进展,也难以实现稳健的长时程操作。尽管分层方法引入了任务分解,但它们主要依赖离线演示的监督学习,无法通过在线交互有效提升执行效果。为解决这一局限,我们提出分层机器人控制框架(Hierarchical Robotic Control, HiRoC),该框架将高层任务规划与低层动作执行解耦。规划器将复杂任务分解为可执行的子目标,以提供显式语义指导;执行器则通过强化学习不断提升子目标条件下的动作生成能力。为使两个模块实现有效协作,我们还在强化学习前将执行器与规划器生成的子目标对齐,以缓解规划与执行间的分布偏差。在多种机器人操作基准上开展的大量实验表明,HiRoC始终优于多个强基线方法。全面分析进一步验证了分层后训练的有效性及各关键组件的贡献。

英文摘要:

Vision-language-action (VLA) models have demonstrated remarkable capabilities in robotic manipulation by leveraging pretrained vision-language models. However, existing post-training methods predominantly optimize VLA models as flat policies, making it difficult to explicitly model task progression and perform robust long-horizon manipulation. Although hierarchical approaches introduce task decomposition, they mainly rely on supervised learning from offline demonstrations and cannot effectively improve execution through online interaction. To address this limitation, we propose Hierarchical Robotic Control (HiRoC), a hierarchical post-training framework that decouples high-level task planning from low-level action execution. The planner decomposes complex tasks into executable subgoals to provide explicit semantic guidance, while the executor continuously improves subgoal-conditioned action generation through reinforcement learning. To enable effective collaboration between the two modules, we further align the executor with planner-generated subgoals before reinforcement learning, mitigating the distribution misalignment between planning and execution. Extensive experiments across diverse robotic manipulation benchmarks demonstrate that HiRoC consistently outperforms strong baselines. Comprehensive analyses further validate the effectiveness of hierarchical post-training and the contribution of each key component.

↑