通过对象中心和几何约束实现抗杂波的视觉-语言-动作模型
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
浏览论文内容
中文总结 AI 辅助
本文提出OBEYED-VLA框架,通过分离感知接地与动作推理提升视觉-语言-动作模型在现实环境中的鲁棒性,特别是在存在干扰物、目标缺失和背景变化等挑战性场景中表现优异。
中文摘要 AI 辅助
近期的视觉-语言-动作(VLA)模型通过在大视觉-语言模型(VLM)上进行微调,实现了在通用机器人操作中的显著进展。然而,大多数VLA模型将感知与控制整合在单一管道中,仅优化动作预测,这会削弱语言条件下的接地能力。在我们的现实桌面上 tabletop 测试中,策略在目标缺失时过度抓取,受到杂乱物干扰,并过度适应背景外观。为了解决这些问题,我们提出了OBEYED-VLA(OBject-centric and gEometrY groundED VLA),一个框架,明确将感知接地与动作推理分离。OBEYED-VLA 不再直接操作原始RGB图像,而是通过一个感知模块将多视角输入转换为任务条件、对象中心和几何感知的观察。该模块包括一个基于VLM的对象中心接地阶段,选择跨摄像头视角的相关物体区域,以及一个互补的几何接地阶段,强调这些物体的3D结构而非外观。生成的接地视图随后输入到预训练的VLA策略中,我们仅在无环境杂乱物或非目标物体的情况下微调单物体演示。在现实世界UR10e桌面设置上,OBEYED-VLA在四个具有挑战性的领域和多个难度级别上显著优于强大的VLA基线:干扰物、目标缺失拒绝、背景外观变化和未见过物体的杂乱操作。消融研究证实,语义接地和几何感知的接地对于这些增益至关重要。总体而言,结果表明,使感知成为显式、对象中心的组件是增强和推广基于VLA的机器人操作的有效方法。
英文摘要
Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action prediction. Yet most VLAs entangle perception and control in a monolithic pipeline optimized purely for action, which can erode language-conditioned grounding. In our real-world tabletop tests, policies over-grasp when the target is absent, are distracted by clutter, and overfit to background appearance. To address these issues, we propose OBEYED-VLA (OBject-centric and gEometrY groundED VLA), a framework that explicitly disentangles perceptual grounding from action reasoning. Instead of operating directly on raw RGB, OBEYED-VLA augments VLAs with a perception module that grounds multi-view inputs into task-conditioned, object-centric, and geometry-aware observations. This module includes a VLM-based object-centric grounding stage that selects task-relevant object regions across camera views, along with a complementary geometric grounding stage that emphasizes the 3D structure of these objects over their appearance. The resulting grounded views are then fed to a pretrained VLA policy, which we fine-tune exclusively on single-object demonstrations collected without environmental clutter or non-target objects. On a real-world UR10e tabletop setup, OBEYED-VLA substantially improves robustness over strong VLA baselines across four challenging regimes and multiple difficulty levels: distractor objects, absent-target rejection, background appearance changes, and cluttered manipulation of unseen objects. Ablation studies confirm that both semantic grounding and geometry-aware grounding are critical to these gains. Overall, the results indicate that making perception an explicit, object-centric component is an effective way to strengthen and generalize VLA-based robotic manipulation.
发表机构
- University of Arkansas(阿拉巴马大学)
- National University of Singapore(新加坡国立大学)
- TU Wien(维也纳技术大学)
- Max Planck Research School for Intelligent Systems(智能系统马克斯·普朗克研究学校)
- University of Stuttgart(斯图加特大学)
- University of Liverpool(利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。