AI 中文总结
GT-VLA提出目标条件化轨迹引导框架,结合通用VLM常识与VLA控制,通过轨迹条件化动作生成提升机器人操作在未见任务和长时程场景中的泛化能力。
AI 中文摘要
视觉-语言-动作(VLA)模型在机器人操作任务中表现出强大性能,但往往难以泛化到未见过的任务、配置和长时程场景。一个关键挑战是VLA模型过度拟合训练场景,无法遵循新颖的语言指令。现成的视觉-语言模型(VLM)通常提供更强的泛化能力,但无法直接控制机器人动作。为了将VLM的常识与VLA的控制能力相结合,我们提出了引导轨迹VLA(GT-VLA),这是一个可引导的框架,通过轨迹条件化的动作生成来接受外部通用VLM的引导。GT-VLA使用通用模型识别当前技能的语义引导,将此引导转换为2D视觉轨迹,并使其动作策略基于生成的轨迹渲染观察进行条件化。该设计将语义目标获取、轨迹生成和低级动作执行分离,使高层引导能够传播到机器人动作中。GT-VLA采用混合专家架构,配备技能特定的轨迹和动作模块以实现稳健执行。我们在LIBERO和物理机器人平台上评估了GT-VLA,在两种设置中均显示出相较于近期VLA基线的改进泛化性能。代码和补充材料可在我们的项目网站上获取,网址为https://this URL。
英文摘要
Vision-Language-Action (VLA) models have shown strong performance on robotic manipulation, but they often struggle to generalize to unseen tasks, configurations, and long-horizon settings. A key challenge is that VLAs overfit to training scenes and fail to follow novel language instructions. Off-the-shelf vision-language models (VLMs) often provide stronger generalization, but cannot directly control robot actions. To combine the common sense of VLMs with VLA control, we propose Guided Trace VLA (GT-VLA), a steerable framework that accepts guidance from an external generalist VLM through trace-conditioned action generation. GT-VLA uses a generalist model to identify semantic guidance for the current skill, converts this guidance into a 2D visual trace, and conditions its action policy on the resulting trace-rendered observation. This design separates semantic target acquisition, trace generation, and low-level action execution, allowing high-level guidance to propagate to robot actions. GT-VLA uses a Mixture-of-Experts architecture with skill-specific trace and action modules for robust execution. We evaluate GT-VLA on LIBERO and a physical robot platform, showing improved generalization over recent VLA baselines in both settings. The code and additional supplemental materials are available on our project website at https://ivaniz.github.io/gt-vla/.