发表机构
Toyota Technological Institute at Chicago; Argonne National Laboratory(芝加哥丰田技术研究所; 阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ProAct动作分词训练方法,通过增强重建目标提升下游可预测性和稳健性,在多个基准和真实场景中显著提高机器人策略成功率。
AI 中文摘要
自回归动作分词策略(如视觉-语言-动作模型)需要动作分词器将离散的标记序列转换为连续空间中的精确控制动作。许多动作分词器通过重建目标来学习标记与动作之间的映射。然而,正如我们通过广泛分析所展示的,足够准确的动作重建只是下游机器人策略成功的一部分。同样关键的是,策略能够为新观察预测正确的标记,并且未见过的策略标记预测仍能解码为合理的动作。这些属性是分词器训练的下游结果,仅靠重建目标并不能直接激励。在这项工作中,我们引入了可预测且稳健的动作分词(ProAct),一种分词器训练方法,策略性地增强重建目标以改善下游可预测性和稳健性。ProAct与策略无关,仅使用动作数据集进行训练。在Robomimic、LIBERO和RoboTwin基准测试以及多种分词器架构中,ProAct将滚动成功率平均提高了11.3个百分点。这些改进也适用于视觉-语言-动作策略和真实世界的机器人操作,分别平均提高了21.8和36.7个百分点。这些结果表明,有效的动作分词应被设计为一种策略接口,平衡保真度、可预测性和稳健性,而不仅仅是重建问题。
英文摘要
Autoregressive action-token policies such as vision-language-action models require action tokenizers to translate discrete token sequences into precise control actions in continuous space. Many action tokenizers learn the mapping between tokens and actions via a reconstruction objective. However, as we show through extensive analysis, sufficiently accurate action reconstruction is only one part of what makes a downstream robot policy successful. It is also critical that the policy is able to predict the right tokens for new observations, and that unseen policy token predictions still decode into reasonable actions. These properties are downstream of tokenizer training and are not directly incentivized by a reconstruction objective alone. In this work, we introduce Predictable and Robust Action Tokenization (ProAct), a tokenizer training method that strategically augments reconstruction with the goal of improving downstream predictability and robustness. ProAct is policy-agnostic and uses only action datasets for training. Across the Robomimic, LIBERO, and RoboTwin benchmarks and a diverse set of tokenizer architectures, ProAct improves rollout success by an average of 11.3 percentage points. These improvements also translate to vision-language-action policies and real-world robotic manipulation, yielding average gains of 21.8 and 36.7 percentage points, respectively. These results suggest that effective action tokenization should be designed as a policy interface that balances fidelity, predictability, and robustness, rather than as a reconstruction problem alone.