发表机构
VinRobotics; Austrian Institute of Technology; Hanoi University of Science and Technology; Center for AI Research, VinUniversity; German Research Center for Artificial Intelligence (DFKI); University of Stuttgart; International Max Planck Research School for Intelligent Systems (IMPRS-IS); University of Arkansas(VinRobotics; 奥地利理工学院; 河内理工大学; VinUniversity人工智能研究中心; 德国人工智能研究中心(DFKI); 斯图加特大学; 国际马克斯·普朗克智能系统研究学院(IMPRS-IS); 阿肯色大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GloVLA将对象中心操作分解为几何传输与局部VLA交互,提升非结构化环境下的鲁棒性和效率,在LIBERO-Challenge和真实机器人上显著提高成功率并降低推理成本。
AI 中文摘要
视觉-语言-动作(VLA)模型在语言条件机器人操作中展现出良好的泛化能力,但在非结构化环境中部署仍具挑战。单一端到端VLA策略必须同时解决末端执行器到任务相关区域的长距离运输以及到达后的短时程、接触丰富的交互。这种表述效率低下且脆弱:微小的视觉偏移、干扰物、杂乱、遮挡或不利的初始夹爪姿态都可能将策略推离其训练时的局部状态分布,导致任务失败。我们提出GloVLA,一种混合框架,明确将对象中心操作分离为两个互补阶段:几何传输控制器将末端执行器移动到交互中心的交接区域,局部VLA策略仅处理短时程交互阶段。GloVLA是模型无关的,可与不同VLA骨干集成,无需额外演示,也无需改变动作空间或成功谓词。在标准LIBERO和LIBERO-Plus Object任务以及新引入的LIBERO-Challenge基准(包含杂乱、干扰物、光照变化、视觉偏移和遮挡)上的实验表明,与全端到端、全轨迹GR00T N1.6相比,GloVLA提高了任务成功率并大幅降低了VLA推理成本。在LIBERO-Challenge中,全轨迹执行平均成功率降至20.9%,而GloVLA保持88.5%;在物理UR10e上,总体成功率从35.6%提升至90.0%,平均推理时间减少一半以上。视频和更多结果可在该https URL获取。
英文摘要
Vision-language-action (VLA) models have shown promising generalization for language-conditioned robot manipulation, but deploying them in unstructured environments remains challenging. A single end-to-end VLA policy must simultaneously solve long-range transport of the end effector to task-relevant regions and short-horizon, contact-rich interaction upon arrival. This formulation is inefficient and brittle: small visual shifts, distractors, clutter, occlusions, or unfavorable initial gripper poses can push the policy outside the local state distribution in which it was trained, leading to task failure. We introduce GloVLA, a hybrid framework that explicitly separates object-centric manipulation into two complementary regimes: a geometric transport controller moves the end-effector into interaction-centric handoff regions, and local VLA policies handle only the short-horizon interaction phases. GloVLA is model-agnostic and can be integrated with different VLA backbones with no additional demonstrations and no changes to the action space or success predicate. Experiments on standard LIBERO and LIBERO-Plus Object tasks together with a newly introduced LIBERO-Challenge benchmark ettings with clutter, distractors,illumination changes, visual shifts, and obstruction show that GloVLA improves task success and substantially lowers VLA inference cost compared with full end-to-Challenge, full-trajectory GR00T N1.6execution degrades to 20.9% average success while GloVLA retains 88.5%; on a physical UR10e, overall success improves from 35.6% to 90.0% while mean inference time is more than halved. Videos and additional results are available at https://glovla-project.github.io/
Comments9 pages, 7 figures. Submitted to IEEE Robotics and Automation Letters (RA-L)