arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09808cs.RO

GTA-2:基于接地任务轴的多VLM机器人操作技能合成框架

GTA-2: A Multi-VLM Framework for Synthesizing Robot Manipulation Skills via Grounded Task Axes

M. Yunus Seker, Shobhit Aggarwal, Ruwan Wickramarachchi, Jonathan Francis, Oliver Kroemer

首次发表
浏览论文内容

中文总结 AI 辅助

GTA-2是一个多VLM框架,通过分解任务轴组件生成零样本机器人操作技能,无需演示或训练,在14个真实任务上平均成功率73.9%,经反馈提升至90.7%。

中文摘要 AI 辅助

机器人操作任务通常被分解为行为或技能。然而,人们往往需要为特定任务预定义这些行为,或者尝试使用通用技能覆盖广泛的任务范围。因此,这些行为可能仍然过于粗糙,无法揭示执行所需的几何、控制和场景相关决策。我们提出了接地任务轴v2(GTA-2),一个模块化的多VLM框架,从可复用的对象中心任务轴组件构建可执行的、任务定制的操作技能。GTA-2不是端到端预测动作或组合固定的任务级原语,而是将每个技能表示为包含任务相关关键点和轴、控制器组合以及场景相关参数的语义子任务。四个专门的VLM代理分别分解任务、构建抽象的任务轴技能、分配控制器参数,并从RGB-D观测中接地所需的视觉特征。这种从抽象到接地的分解使得无需任务特定的机器人演示、策略训练或微调即可实现零样本技能生成。它还保持中间决策的明确性,允许有针对性的反馈来改进不正确的阶段,同时保留正确的组件。我们在14个真实机器人操作任务上评估了GTA-2,与VLA策略pi_{0.5}和两个使用任务轴控制器或传统机器人原语的Code-as-Policies基线进行比较。GTA-2的平均零样本成功率达到73.9%,超过最强基线31.4个百分点,而有针对性的改进将GTA-2的平均成功率提高到90.7%。项目页面:此https URL

英文摘要

Robotic manipulation tasks are often decomposed into behaviors or skills. However, one often needs to predefine these behaviors for specific tasks or try to cover a wide range of tasks using generic skills. As a result, these behaviors can remain too coarse to expose the geometric, control, and scene-dependent decisions required for execution. We introduce Grounded Task Axes v2 (GTA-2), a modular multi-VLM framework that constructs executable, task-bespoke manipulation skills from reusable object-centric task-axis components. Rather than predicting actions end-to-end or composing fixed task-level primitives, GTA-2 represents each skill as semantic subtasks comprising task-relevant keypoints and axes, controller compositions, and scene-dependent parameters. Four specialized VLM agents separately decompose the task, construct an abstract task-axis skill, assign controller parameters, and ground the required visual features from RGB-D observations. This abstraction-to-grounding factorization enables zero-shot skill generation without task-specific robot demonstrations, policy training, or fine-tuning. It also keeps intermediate decisions explicit, allowing targeted human feedback to refine an incorrect stage while preserving correct components. We evaluate GTA-2 on 14 real-robot manipulation tasks against a VLA policy pi_{0.5} and two Code-as-Policies baselines using task-axis controllers or conventional robot primitives. GTA-2 achieves an average zero-shot success rate of 73.9%, exceeding the strongest baseline by 31.4 percentage points, while targeted refinement raises GTA-2's average success rate to 90.7%. Project page: https://gta2-project.github.io/

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • Bosch Research(博世研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑