BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
BridgeVLA++:一种面向三维操作的数据高效、可泛化且内存增强的视觉-语言-动作框架
机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室(NLPR)) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; FiveAges ; ByteDance Seed(字节跳动种子实验室)
AI总结 本研究提出内存增强的三维VLA框架BridgeVLA++,通过新增时空记忆架构,在保留原模型数据效率与泛化能力的同时,提升了记忆相关操作性能,且在多任务与真实平台上验证了其有效性。
Comments This work has been submitted to the IEEE TPAMI for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible