arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

${M}^2$Tok:用于视觉-语言-动作模型的多头多码本离散动作标记化

M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

Chunpu Xu, Zhixuan Liang, Yuhao Zhang, Chi-Min Chan, Jessie Wang, Yang Xiao, Mengkang Hu, Xiaokang Yang, Yao Mu

arXiv 2609.18259首次发表:更新:

发表机构

The Hong Kong Polytechnic University; Shanghai AI Laboratory; The University of Hong Kong; Shanghai Jiao Tong University; Hong Kong University of Science and Technology(香港理工大学; 上海人工智能实验室; 香港大学; 上海交通大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出多头多码本动作标记器M2Tok,通过分解动作特征并分配独立码本,降低重建损失,提升VLA模型在仿真和真实任务中的成功率。

AI 中文摘要

近期研究进展已成功将自回归语言模型适配于处理多模态信号,如图像和动作。由于原始动作信号是连续的,有效的标记化对于将高维输入映射为紧凑的离散标记以进行自回归处理至关重要。然而,现有的离散动作标记器常常遭受高重建损失,无法保留精确控制所需的细粒度动态信息。这种“离散化瓶颈”显著限制了下游视觉-语言-动作(VLA)模型的性能上限。为解决此问题,我们提出了$\mathcal{M}^2$Tok,一种多头多码本动作标记器,旨在最小化重建误差并提升策略性能。我们的方法引入了两项关键结构创新:(1)我们将潜在动作特征分解为多个头,使模型能够隐式地将特定头与不同动作维度对齐;(2)我们为每个头分配独立的码本进行量化。通过利用多个码本的组合特性,我们显著扩展了标记器的表示表达能力,与先前方法相比,大幅降低了重建损失。我们在RoboTwin、Simpler-Env以及3个零样本真实世界任务上评估了基于$\mathcal{M}^2$Tok的VLA模型。实验结果表明,我们的方法不仅实现了卓越的重建保真度,还显著提升了VLA模型的成功率。全面的消融研究进一步证实了多头和多码本机制的有效性。代码可在\href{this https URL}{this https URL}获取。

英文摘要

Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This "discretization bottleneck" significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose ${M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the ${M}^2$Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at https://github.com/cpaaax/M2Tok.

CommentsECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑