arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BridgeVLA++:一种面向三维操作的数据高效、可泛化且内存增强的视觉-语言-动作框架

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma, Yuan Xu, Jiabing Yang, He Guan, Yan Huang, Hongtao Wu, Xiao Ma, Tao Kong, Liang Wang, Tieniu Tan

arXiv 2608.05042首次发表:更新:

发表机构

New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; FiveAges; ByteDance Seed(中国科学院自动化研究所模式识别国家重点实验室(NLPR); 中国科学院大学人工智能学院; FiveAges; 字节跳动种子实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出内存增强的三维VLA框架BridgeVLA++,通过新增时空记忆架构,在保留原模型数据效率与泛化能力的同时,提升了记忆相关操作性能,且在多任务与真实平台上验证了其有效性。

AI 中文摘要

利用预训练视觉-语言模型(VLMs)构建视觉-语言-动作(VLA)模型已成为三维机器人操作领域极具前景的范式。然而,现有的三维VLA方法仍存在数据需求量大、分布偏移下泛化能力有限、缺乏对过往观测的显式记忆等问题,这些局限性阻碍了其在数据稀缺、开放世界及依赖记忆的操作场景中的应用。我们之前的工作BridgeVLA通过在三维动作学习过程中保留预训练VLM的输入-输出对齐,提升了数据效率与泛化能力:将原始点云投影为多视角图像,在生成机器人动作前预测中间热图。本研究中,我们为BridgeVLA配备了统一的时空记忆架构,该架构可对持久的空间上下文与时间交互历史进行建模,由此得到的内存增强框架能够在保留BridgeVLA数据效率与泛化能力的同时,对观测历史进行推理。大量实验表明,该框架在空间操作任务上表现出色,且具有较强的泛化能力;BridgeVLA++在两个具有挑战性的依赖记忆的操作基准测试中达到了当前最优性能,且未牺牲原始BridgeVLA的数据效率与泛化能力。此外,BridgeVLA++在双臂操作场景中表现良好,并在另一个真实世界机器人平台上得到验证,证明了其在任务、环境及机器人平台间的可扩展性。这些结果确立了BridgeVLA++是一种统一的三维视觉-语言-动作框架,可同时支持数据高效学习、鲁棒泛化及有效的记忆感知机器人操作。项目网站:this https URL

英文摘要

Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and memory-dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input--output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real-world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization, and effective memory-aware robot manipulation. Project website: https://bridgevla-plus.github.io/.

CommentsThis work has been submitted to the IEEE TPAMI for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑