arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02257cs.RO

面向全身遥操作移动操作的全景感知VLA学习

Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation

  • Beihang University(北京航空航天大学)
  • Shandong University(山东大学)
  • Johns Hopkins University(约翰斯·霍普金斯大学)
  • Delta Intelligence(delta智能公司)

机构由 AI 辅助整理,请以论文原文为准。

Donglin Yang, Haoran Chen, Xingyu Chen, Lixing Liu, Manyi Li, Changhe Tu, Ke Xu, Xiaojian Ma, Si Liu

AI总结:

针对移动操作VLA策略的局部视野与全身演示数据收集难题,开发全身遥操作系统与PanoVLA全景感知VLA策略,基于5.5小时多模态数据集,在四项真实任务中平均阶段完成率91.3%、端到端成功率73.4%,性能优于局部视角基线。

AI中文摘要:

移动操作是具身智能的核心能力,使机器人能在开放世界环境中完成复杂多阶段任务。然而移动操作对视觉-语言-动作(VLA)策略提出两大关键挑战:数据层面,高效收集高质量全身演示需协调控制移动基座与机械臂;模型层面,现有VLA模型主要依赖局部相机观测,其有限视野阻碍全局空间理解。为应对这些挑战,我们开发了全身遥操作系统与全景感知VLA策略。该系统通过单个VR界面实现轮式双臂机器人的协调控制,支持采集包含5.5小时多模态演示的真实世界移动操作数据集。基于此数据集,我们提出PanoVLA,一种面向移动双臂操作的全景感知视觉-语言-动作策略。PanoVLA基于混合Transformer架构,通过专用全景编码与融合模块引入全局空间上下文,能有效将全景观测与语言指令、机器人状态整合以生成动作。在四项真实世界移动操作任务上的评估显示,PanoVLA的平均阶段完成率达91.3%,端到端成功率为73.4%,显著优于局部视角基线。这些结果表明,融入全景空间上下文可提升移动机器人的空间理解与闭环操作性能。

英文摘要:

Mobile manipulation is a key capability for embodied intelligence, enabling robots to accomplish complex multi-stage tasks in open-world environments. However, mobile manipulation poses two key challenges for vision-language-action (VLA) policies: At the data level, the efficient collection of high-quality whole-body demonstrations demands the coordinated control of both the mobile base and the robotic arms; at the model level, existing VLA models predominantly rely on local camera observations, whose limited field of view hinders global spatial understanding. To address these challenges, we develop a whole-body teleoperation system and a panoramic-aware VLA policy. The system enables coordinated control of a wheeled bimanual robot through a single VR interface and supports the acquisition of a real-world mobile manipulation dataset comprising 5.5 hours of multimodal demonstrations. Building upon this dataset, we propose PanoVLA, a panorama-aware vision-language-action policy for mobile bimanual manipulation. Built upon a Mixture-of-Transformers architecture, PanoVLA introduces global spatial context through dedicated panorama encoding and fusion modules, enabling effective integration of panoramic observations with language instructions and robot states for action generation. Evaluation on four real-world mobile manipulation tasks demonstrates that PanoVLA achieves an average stage completion rate of 91.3\% and an end-to-end success rate of 73.4\%, substantially outperforming local-view baselines. These results demonstrate that incorporating panoramic spatial context improves spatial understanding and closed-loop manipulation performance in mobile robots.

补充信息

↑