AI 中文总结
研究针对动作令牌化存在的问题,提出有序动作令牌化(OAT)方法,利用带特定机制的变换器将动作块离散化,满足高压缩率等三个需求,在多种策略和任务中验证,能提升策略性能并增强推理灵活性。
AI 中文摘要
动作令牌化将连续的机器人动作块映射到离散令牌,已成为现代视觉运动策略的重要接口。现有方法要么依赖于产生过长令牌序列的解析离散化方法,要么依赖于缺乏结构的学习潜在令牌化器,限制了它们与下游策略的兼容性。在这项工作中,我们确定了动作令牌化的三个需求——高压缩率、完全可解码性和有序令牌空间,并引入了有序动作令牌化(OAT),一种满足所有这三个需求的学习动作令牌化器。OAT使用带有寄存器、有限标量量化和排序诱导训练机制的变换器将动作块离散化为有序的令牌序列。通过训练每个令牌前缀解码为有效的动作块,OAT将粗略控制信息置于早期令牌中,并使用后期令牌细化残余细节,在推理成本和动作保真度之间产生随时权衡。我们在动作令牌的两种主要用途中验证了OAT:为控制生成令牌的自回归策略,以及使用令牌损失来塑造基于流的动作专家消耗的视觉语言模型上下文的令牌协同训练策略。在跨越五个模拟基准和现实世界设置的三个策略主干和60多个任务中,OAT始终提供强大的策略性能,同时在推理时提供显著更大的灵活性。
英文摘要
Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on analytical discretization methods that produce prohibitively long token sequences or learned latent tokenizers that lack structure, limiting their compatibility with downstream policies. In this work, we identify three desiderata for action tokenization - high compression, total decodability, and an ordered token space - and introduce Ordered Action Tokenization (OAT), a learned action tokenizer that satisfies all three. OAT discretizes action chunks into an ordered sequence of tokens using a transformer with registers, finite scalar quantization, and ordering-inducing training mechanisms. By training each token prefix to decode into a valid action chunk, OAT places coarse control information in early tokens and uses later tokens to refine residual detail, yielding an anytime tradeoff between inference cost and action fidelity. We validate OAT in two prevailing uses of action tokens: autoregressive policies that generate tokens for control, and token co-training policies that use token losses to shape the vision-language model context consumed by a flow-based action expert. Across three policy backbones and more than 60 tasks spanning five simulation benchmarks and real-world settings, OAT consistently delivers strong policy performance while offering significantly greater flexibility at inference time.