发表机构
Nanyang Technological University; ACE Robotics(南洋理工大学; ACE机器人公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ACT³模型,通过分层注意力融合VLM与WM的语义和动力学信息,在模拟及真实机器人操作基准上取得优于同类方法的结果,提升了机器人策略泛化能力。
AI 中文摘要
视觉-语言-动作(VLA)模型已成为复杂机器人操作领域的重要框架,它基于预训练视觉-语言模型(VLM)强大的语义理解能力构建。然而,这类VLM主干网络提供的物理动力学先验不足,限制了机器人策略的泛化能力。因此,近期研究尝试通过多种策略将视频生成世界模型(WM)集成到机器人策略中,利用预测动力学辅助动作生成。尽管取得了这些进展,但将语义理解与动力学预测作为互补指导用于动作生成仍具挑战性。本文提出了ACT³,一种简单而有效的以动作中心的三流Transformer,它将语义和动力学信息融合到控制动作中,同时保留各上下文流的独特作用。具体而言,ACT³使专用动作专家通过分层注意力访问VLM和WM的表示,每个主干仅在自身流内进行注意力计算。这种简洁的交互设计保持了上下文流的独立前向传播,同时允许两个主干通过控制监督进行更新。在模拟和真实世界机器人操作基准上的实验表明,所提出的ACT³取得了优于同类方法的结果。
英文摘要
Vision-Language-Action (VLA) models have emerged as a prominent framework for complex robotic manipulation, building on the strong semantic understanding of pretrained Vision-Language Models (VLMs). However, such VLM backbones offer insufficient physical dynamics priors, which limits the generalization capabilities of robot policies. Recent efforts therefore integrate video-generation World Models (WMs) into robot policies through various strategies, using predictive dynamics to facilitate action generation. Despite these advances, harnessing semantic understanding and dynamics prediction as complementary guidance for action generation remains challenging. In this paper, we introduce $\mathrm{ACT}^3$, a simple yet effective Action-Centric Tri-Stream Transformer that fuses semantic and dynamics information into control actions while preserving the distinct roles of context streams. Specifically, $\mathrm{ACT}^3$ enables the dedicated action expert to access VLM and WM representations through layerwise attention, with each backbone attending only within its own stream. This straightforward interaction design maintains independent forward propagation in the context streams while allowing both backbones to be updated through control supervision. Experiments on both simulated and real-world robotic manipulation benchmarks show that the proposed $\mathrm{ACT}^3$ yields results superior to its counterparts.