arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16837cs.ROcs.AI

HAF:通过分层动作流与谱潜在线性RL将通用VLAs适配至人形机器人全身运动操作

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang, Jiaming Liu, Xinyuan Lv, Bowei Zhang, Shuanghao Bai, Guangrun Li, Jingyang He, Gaole Dai, Ziluo Ding, Zhiyuan Xu… 展开作者

Langzhe Gu, Chengkai Hou, Meng Li, Xinhua Wang, Jiaming Liu, Xinyuan Lv, Bowei Zhang, Shuanghao Bai, Guangrun Li, Jingyang He, Gaole Dai, Ziluo Ding, Zhiyuan Xu, Kuan Cheng, Jian Tang, Zhengping Che, Shanghang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对通用VLAs难以适配人形机器人全身运动操作的问题,提出HAF框架,通过分层动作流生成器与潜空间RL流水线实现高效迁移优化,在7项任务上性能优于单阶段VLA基线。

中文摘要 AI 辅助

人形机器人在以人为中心的环境中作为通用智能体具有巨大潜力,但通用视觉-语言-动作(VLA)基础模型难以直接应用于人形机器人全身运动操作。人形机器人运动的高维性与相互依赖性,使得传统单阶段VLA架构难以有效协调 locomotion( locomotion 译为运动)、腰部姿态及双臂操作。此外,通过离线行为克隆训练的策略在实际部署时可能仍处于次优状态。尽管在线强化学习可通过实际交互优化策略,但直接调优大型VLA主干网络需要过多计算资源,且在实际机器人探索过程中可能引入安全风险。为解决这些瓶颈,我们提出HAF(Humanoid Adaptation Framework,人形机器人适配框架),这是一个由HAF-VLA和HAF-Steer组成的两部分框架,用于将现成的通用VLA基础模型迁移至人形机器人全身运动操作。HAF-VLA是构建在预训练的流匹配VLA之上的分层动作流生成器,它将全身动作去噪拆分为三个连续阶段,并利用阶段嵌入与跨阶段KV缓存保留运动学依赖关系,避免单次生成导致的全身动作不连贯。在冻结的HAF-VLA之上,HAF-Steer是一个潜空间离线到在线的RL流水线,它利用流匹配的可逆性与基于DCT的降维技术,将RL优化限制在紧凑的噪声子空间中,并训练一个正则化的SAC策略。这避免了对大型VLA主干网络的更新,实现了高效的实际机器人策略优化。在7项实际人形机器人运动操作任务上的评估表明,HAF优于普通单阶段VLA基线,提升了全身协调性与任务性能。项目网站:this https URL

英文摘要

Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we introduce HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation. HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA. It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches that retain kinematic dependencies, avoiding incoherent whole-body actions from one-shot generation. On top of the frozen HAF-VLA, HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction to restrict RL optimization to a compact noise subspace and train a regularized SAC policy. This avoids updating the large VLA backbone and enables efficient real-world policy refinement. Evaluated on seven real-world humanoid loco-manipulation tasks, HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance. Project website: https://grange007.github.io/HAF .

发表机构

  • Xi’an Jiaotong University(西安交通大学)
  • Peking University(北京大学)
  • Nankai University(南开大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑