HAF:通过分层动作流与谱潜在线性RL将通用VLAs适配至人形机器人全身运动操作
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL
浏览论文内容
中文总结 AI 辅助
研究针对通用VLAs难以适配人形机器人全身运动操作的问题,提出HAF框架,通过分层动作流生成器与潜空间RL流水线实现高效迁移优化,在7项任务上性能优于单阶段VLA基线。
中文摘要 AI 辅助
人形机器人在以人为中心的环境中作为通用智能体具有巨大潜力,但通用视觉-语言-动作(VLA)基础模型难以直接应用于人形机器人全身运动操作。人形机器人运动的高维性与相互依赖性,使得传统单阶段VLA架构难以有效协调 locomotion( locomotion 译为运动)、腰部姿态及双臂操作。此外,通过离线行为克隆训练的策略在实际部署时可能仍处于次优状态。尽管在线强化学习可通过实际交互优化策略,但直接调优大型VLA主干网络需要过多计算资源,且在实际机器人探索过程中可能引入安全风险。为解决这些瓶颈,我们提出HAF(Humanoid Adaptation Framework,人形机器人适配框架),这是一个由HAF-VLA和HAF-Steer组成的两部分框架,用于将现成的通用VLA基础模型迁移至人形机器人全身运动操作。HAF-VLA是构建在预训练的流匹配VLA之上的分层动作流生成器,它将全身动作去噪拆分为三个连续阶段,并利用阶段嵌入与跨阶段KV缓存保留运动学依赖关系,避免单次生成导致的全身动作不连贯。在冻结的HAF-VLA之上,HAF-Steer是一个潜空间离线到在线的RL流水线,它利用流匹配的可逆性与基于DCT的降维技术,将RL优化限制在紧凑的噪声子空间中,并训练一个正则化的SAC策略。这避免了对大型VLA主干网络的更新,实现了高效的实际机器人策略优化。在7项实际人形机器人运动操作任务上的评估表明,HAF优于普通单阶段VLA基线,提升了全身协调性与任务性能。项目网站:this https URL
英文摘要
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we introduce HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation. HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA. It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches that retain kinematic dependencies, avoiding incoherent whole-body actions from one-shot generation. On top of the frozen HAF-VLA, HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction to restrict RL optimization to a compact noise subspace and train a regularized SAC policy. This avoids updating the large VLA backbone and enables efficient real-world policy refinement. Evaluated on seven real-world humanoid loco-manipulation tasks, HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance. Project website: https://grange007.github.io/HAF .
发表机构
- Xi’an Jiaotong University(西安交通大学)
- Peking University(北京大学)
- Nankai University(南开大学)
机构由 AI 辅助整理,请以论文原文为准。