GTA-2:基于接地任务轴的多VLM机器人操作技能合成框架
GTA-2: A Multi-VLM Framework for Synthesizing Robot Manipulation Skills via Grounded Task Axes
浏览论文内容
中文总结 AI 辅助
GTA-2是一个多VLM框架,通过分解任务轴组件生成零样本机器人操作技能,无需演示或训练,在14个真实任务上平均成功率73.9%,经反馈提升至90.7%。
中文摘要 AI 辅助
机器人操作任务通常被分解为行为或技能。然而,人们往往需要为特定任务预定义这些行为,或者尝试使用通用技能覆盖广泛的任务范围。因此,这些行为可能仍然过于粗糙,无法揭示执行所需的几何、控制和场景相关决策。我们提出了接地任务轴v2(GTA-2),一个模块化的多VLM框架,从可复用的对象中心任务轴组件构建可执行的、任务定制的操作技能。GTA-2不是端到端预测动作或组合固定的任务级原语,而是将每个技能表示为包含任务相关关键点和轴、控制器组合以及场景相关参数的语义子任务。四个专门的VLM代理分别分解任务、构建抽象的任务轴技能、分配控制器参数,并从RGB-D观测中接地所需的视觉特征。这种从抽象到接地的分解使得无需任务特定的机器人演示、策略训练或微调即可实现零样本技能生成。它还保持中间决策的明确性,允许有针对性的反馈来改进不正确的阶段,同时保留正确的组件。我们在14个真实机器人操作任务上评估了GTA-2,与VLA策略pi_{0.5}和两个使用任务轴控制器或传统机器人原语的Code-as-Policies基线进行比较。GTA-2的平均零样本成功率达到73.9%,超过最强基线31.4个百分点,而有针对性的改进将GTA-2的平均成功率提高到90.7%。项目页面:此https URL
英文摘要
Robotic manipulation tasks are often decomposed into behaviors or skills. However, one often needs to predefine these behaviors for specific tasks or try to cover a wide range of tasks using generic skills. As a result, these behaviors can remain too coarse to expose the geometric, control, and scene-dependent decisions required for execution. We introduce Grounded Task Axes v2 (GTA-2), a modular multi-VLM framework that constructs executable, task-bespoke manipulation skills from reusable object-centric task-axis components. Rather than predicting actions end-to-end or composing fixed task-level primitives, GTA-2 represents each skill as semantic subtasks comprising task-relevant keypoints and axes, controller compositions, and scene-dependent parameters. Four specialized VLM agents separately decompose the task, construct an abstract task-axis skill, assign controller parameters, and ground the required visual features from RGB-D observations. This abstraction-to-grounding factorization enables zero-shot skill generation without task-specific robot demonstrations, policy training, or fine-tuning. It also keeps intermediate decisions explicit, allowing targeted human feedback to refine an incorrect stage while preserving correct components. We evaluate GTA-2 on 14 real-robot manipulation tasks against a VLA policy pi_{0.5} and two Code-as-Policies baselines using task-axis controllers or conventional robot primitives. GTA-2 achieves an average zero-shot success rate of 73.9%, exceeding the strongest baseline by 31.4 percentage points, while targeted refinement raises GTA-2's average success rate to 90.7%. Project page: https://gta2-project.github.io/
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
- Bosch Research(博世研究院)
机构由 AI 辅助整理,请以论文原文为准。