arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.22519cs.RO

通过对象中心和几何约束实现抗杂波的视觉-语言-动作模型

Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding

Khoa Vo, Taisei Hanyu, Yuki Ikebe, Trong Thang Pham, Nhat Chung, Minh Nhat Vu, Duy Nguyen Ho Minh, Anh Nguyen, Anthony Gunderman, Chase Rainwater, Ngan Le

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出OBEYED-VLA框架,通过分离感知接地与动作推理提升视觉-语言-动作模型在现实环境中的鲁棒性,特别是在存在干扰物、目标缺失和背景变化等挑战性场景中表现优异。

中文摘要 AI 辅助

近期的视觉-语言-动作(VLA)模型通过在大视觉-语言模型(VLM)上进行微调,实现了在通用机器人操作中的显著进展。然而,大多数VLA模型将感知与控制整合在单一管道中,仅优化动作预测,这会削弱语言条件下的接地能力。在我们的现实桌面上 tabletop 测试中,策略在目标缺失时过度抓取,受到杂乱物干扰,并过度适应背景外观。为了解决这些问题,我们提出了OBEYED-VLA(OBject-centric and gEometrY groundED VLA),一个框架,明确将感知接地与动作推理分离。OBEYED-VLA 不再直接操作原始RGB图像,而是通过一个感知模块将多视角输入转换为任务条件、对象中心和几何感知的观察。该模块包括一个基于VLM的对象中心接地阶段,选择跨摄像头视角的相关物体区域,以及一个互补的几何接地阶段,强调这些物体的3D结构而非外观。生成的接地视图随后输入到预训练的VLA策略中,我们仅在无环境杂乱物或非目标物体的情况下微调单物体演示。在现实世界UR10e桌面设置上,OBEYED-VLA在四个具有挑战性的领域和多个难度级别上显著优于强大的VLA基线:干扰物、目标缺失拒绝、背景外观变化和未见过物体的杂乱操作。消融研究证实,语义接地和几何感知的接地对于这些增益至关重要。总体而言,结果表明,使感知成为显式、对象中心的组件是增强和推广基于VLA的机器人操作的有效方法。

英文摘要

Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action prediction. Yet most VLAs entangle perception and control in a monolithic pipeline optimized purely for action, which can erode language-conditioned grounding. In our real-world tabletop tests, policies over-grasp when the target is absent, are distracted by clutter, and overfit to background appearance. To address these issues, we propose OBEYED-VLA (OBject-centric and gEometrY groundED VLA), a framework that explicitly disentangles perceptual grounding from action reasoning. Instead of operating directly on raw RGB, OBEYED-VLA augments VLAs with a perception module that grounds multi-view inputs into task-conditioned, object-centric, and geometry-aware observations. This module includes a VLM-based object-centric grounding stage that selects task-relevant object regions across camera views, along with a complementary geometric grounding stage that emphasizes the 3D structure of these objects over their appearance. The resulting grounded views are then fed to a pretrained VLA policy, which we fine-tune exclusively on single-object demonstrations collected without environmental clutter or non-target objects. On a real-world UR10e tabletop setup, OBEYED-VLA substantially improves robustness over strong VLA baselines across four challenging regimes and multiple difficulty levels: distractor objects, absent-target rejection, background appearance changes, and cluttered manipulation of unseen objects. Ablation studies confirm that both semantic grounding and geometry-aware grounding are critical to these gains. Overall, the results indicate that making perception an explicit, object-centric component is an effective way to strengthen and generalize VLA-based robotic manipulation.

发表机构

  • University of Arkansas(阿拉巴马大学)
  • National University of Singapore(新加坡国立大学)
  • TU Wien(维也纳技术大学)
  • Max Planck Research School for Intelligent Systems(智能系统马克斯·普朗克研究学校)
  • University of Stuttgart(斯图加特大学)
  • University of Liverpool(利物浦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑