发表机构
Institute of Cyber-Systems and Control, Zhejiang University; Tongji University(浙江大学网络系统与控制研究所; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对视觉-语言-动作模型问题,提出TS-Mask VLA框架,基于离散扩散动作专家和时空二维掩码策略,在模拟基准和现实任务实验中表现出色,验证了设计有效性。
AI 中文摘要
视觉-语言-动作(VLA)模型旨在理解自然语言指令和视觉观察,并作为具身智能体生成和执行相应动作。基于自回归令牌的动作生成推动了许多代表性VLA模型的发展,但存在缺乏动作序列时空结构显式建模等问题。本文提出TS-Mask VLA框架,基于两个关键设计:配备桥接注意力调节桥的离散扩散动作专家,以及离散动作令牌的时空二维掩码策略。在模拟基准和现实任务上进行了大量实验,结果表明该模型在参数较少时仍有出色表现,综合分析和消融实验进一步验证了设计的有效性。
英文摘要
Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. Recently, autoregressive token-based action generation has driven the development of many representative VLA models. However, this paradigm often reduces action generation to next-token prediction, thereby lacking explicit modeling of the spatiotemporal structure of action sequences and the disentanglement between vision-language representations and actions, which can limit performance in long-horizon and complex scenarios. In this paper, we propose TS-Mask VLA, a vision-language-action framework for robot manipulation. TS-Mask VLA is built upon two key designs: (1) a Discrete Diffusion Action Expert equipped with a Bridge Attention conditioning bridge, which enables multi-layer conditioning from the VLM and facilitates more accurate and stable action generation; and (2) a temporal-spatial 2D masking strategy for discrete action tokens that strengthens the model's understanding of cross-time dependencies and inter-dimensional coupling, leading to more structurally consistent action sequences. We conduct extensive experiments on simulation benchmarks and real-world tasks. On LIBERO, TS-Mask VLA achieves a 95.7 percent average success rate with only 0.5B parameters, outperforming significantly larger models. On CALVIN, it attains the best average sequence length of 4.19 and strong long-horizon performance. Comprehensive analyses and ablations further validate the effectiveness of our design.
Comments9 pages, 5 figures, accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)