arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

H-VLA:统一动作空间下具有关键动作推理与运动规划的分层视觉-语言-动作模型

H-VLA: Hierarchical Vision-Language-Action Model with Key-Action Reasoning and Motion Planning in a Unified Action Space

Xiongfeng Peng, Lu Xu, Yandong Wang, Jiaqian Yu, Zirui Zheng, Yamin Mao, Weiming Li, Inseop Chung, Hyun-woong Cho, Jaewook Yoo, Dongwook Lee, Daehyun Ji, Chao Zhang

arXiv 2609.22895首次发表:更新:

发表机构

Advanced Research Lab, Samsung R&D Institute China-Beijing (SRCB); Samsung AI Center, DS Division(三星电子中国研究院北京先进研究实验室; 三星AI中心DS部门)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出分层VLA框架H-VLA,解耦关键动作推理与运动规划,采用统一相机中心动作空间,在SimplerEnv和真实机器人任务上显著提升操作性能。

AI 中文摘要

视觉-语言-动作(VLA)模型在机器人操作中展现出巨大潜力,但许多现有方法仍依赖从语言和视觉观测到密集动作的直接映射。这种表述会削弱从预训练视觉-语言模型(VLM)继承的语义推理能力,这些模型主要针对视觉-语言理解而非低级控制进行优化,并且在空间变化(包括物体位置、场景布局、机器人本体和相机视角的变化)下变得脆弱。为解决这些局限性,我们提出H-VLA,一种分层VLA框架,将高层关键动作推理与低层运动生成解耦。H-VLA结合了用于预测关键动作作为下一个操作子目标的关键动作模型、用于在预测关键动作条件下生成密集未来动作的运动规划模型,以及用于跨数据集、本体和视角一致表示的统一相机中心动作空间。我们进一步采用两阶段训练策略,在预训练阶段强调关键动作推理,在微调阶段强调密集运动生成。实验表明,H-VLA在SimplerEnv上取得了强劲性能,在Google Robot视觉匹配上达到91%,在Google Robot变体聚合上达到84%,在WidowX视觉匹配上达到81%。在Agilex真实机器人任务中,H-VLA在分布内、分布外位置和分布外场景/物体设置下,分别比最强基线提高了10、47和16个百分点。

英文摘要

Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but many existing methods still rely on direct mappings from language and visual observations to dense actions. This formulation can weaken the semantic reasoning capability inherited from pre-trained Vision-Language Models (VLMs), which are mainly optimized for visual-linguistic understanding rather than low-level control, and becomes fragile under spatial variations, including changes in object positions, scene layouts, robot embodiments, and camera viewpoints. To address these limitations, we propose H-VLA, a hierarchical VLA framework that decouples high-level key-action reasoning from low-level motion generation. H-VLA combines a Key-Action Model for predicting a key-action as the next manipulation subgoal, a Motion Planning Model for generating dense future actions conditioned on the predicted key-action, and a Unified Camera-Centric Action Space for consistent representation across datasets, embodiments, and viewpoints. We further adopt a two-stage training strategy that emphasizes key-action reasoning during pre-training and dense motion generation during fine-tuning. Experiments show that H-VLA achieves strong performance on SimplerEnv, reaching 91% on Google Robot visual matching, 84% on Google Robot variant aggregation, and 81% on WidowX visual matching. On Agilex real-robot tasks, H-VLA improves over the strongest baseline by 10, 47, and 16 percentage points under in-distribution, out-of-distribution position, and out-of-distribution scene/object settings, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑