发表机构
Huazhong University of Science and Technology; Zhongguancun Academy; DeepCybo; Harbin Institute of Technology; Beihang University; The Hong Kong University of Science and Technology (Guangzhou); Zhengzhou University; Zhongguancun Institute of Artificial Intelligence(华中科技大学; 中关村学院; DeepCybo; 哈尔滨工业大学; 北京航空航天大学; 香港科技大学(广州); 郑州大学; 中关村人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对自回归VLA模型动作分词保真度不足问题,提出物理秩一致性指标和ActionPiece方法,通过联合监督表示学习与量化保留动作关系,在多个基准上显著提升策略成功率。
AI 中文摘要
动作分词器在自回归视觉-语言-动作(VLA)模型中扮演核心角色,既决定了策略训练的目标,也决定了从预测词元中恢复的可执行命令。其保真度通常使用逐点重建指标(如均方误差(MSE))进行评估,然而小的个体误差并不能完全刻画跨演示的动作调整被保留的忠实程度。压缩后,相似动作可能仍聚集在代表性运动周围,而不同情境所需的调整则被削弱、扭曲甚至反转。我们引入物理秩一致性(PRC)来衡量分词在重建后保留局部物理距离排序的程度。评估解码后的动作可为不同词元词汇表和解码器架构提供共同参考,以关系保真度指标补充逐点精度。我们进一步提出ActionPiece,通过表示学习和量化的联合监督来保留物理动作关系。物理秩保留监督编码器和量化特征距离中的远近排序,而量化正则化则将相同排序应用于码字分配分布。这两个目标增强重建,产生离散动作词元,用于标准自回归策略学习,并通过冻结解码器执行。在相同的Qwen3-VL-4B策略训练设置下,ActionPiece在LIBERO上达到94.8%,在未见过的LIBERO-Plus上达到68.8%,额外评估在SimplerEnv上达到71.9%,在VLA-Arena L0-L2上平均达到51.5%。组件消融表明,这两个目标联合改善了PRC和策略成功率,展示了物理关系监督对动作分词的价值。
英文摘要
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens. Their fidelity is commonly evaluated using pointwise reconstruction metrics such as mean squared error (MSE), yet small individual errors do not fully characterize how faithfully action adjustments across demonstrations are preserved. After compression, similar actions may still cluster around a representative motion, while the adjustments needed for different contexts are diminished, distorted, or even reversed. We introduce physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction. Evaluating decoded actions provides a common reference across token vocabularies and decoder architectures, complementing pointwise accuracy with a measure of relational fidelity. We further present ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization. Physical rank preservation supervises near-far ordering in encoder and quantized feature distances, while quantization regularization applies the same ordering to codeword assignment distributions. Both objectives augment reconstruction, producing discrete action tokens for standard autoregressive policy learning and execution through a frozen decoder. Under the same Qwen3-VL-4B policy training setup, ActionPiece achieves 94.8% on LIBERO and 68.8% on unseen LIBERO-Plus, with additional evaluations reaching 71.9% on SimplerEnv and 51.5% across VLA-Arena L0-L2. Component ablations show that the two objectives jointly improve PRC and policy success, demonstrating the value of physical relationship supervision for action tokenization.
CommentsProject Page: https://deepcybo-physai.github.io/ActionPiece/