发表机构
Toyota Technological Institute at Chicago; Argonne National Laboratory(芝加哥丰田理工学院; 阿贡国家实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对分词化机器人策略忽略分词器潜在码结构的问题,提出码本对齐预测方法,重用码向量作为类别原型,在多个基准和真实任务上稳定提升成功率。
AI 中文摘要
动作分词将连续的机器人动作转换为离散符号,从而可以进行自回归建模。然而,现有的基于分词器的策略通常忽略分词器学习到的潜在码结构:在分词后,策略将分词视为无关的类别索引,并从头学习一个新的分类器。我们表明,这种被丢弃的结构是有价值的。我们引入了码本对齐预测(CAP),一种直接重用分词器的码向量作为策略类别原型的方法,同时保持分词器和策略主干不变。在四种量化器家族、三个模拟基准和两个真实机器人任务中,与标准分词分类头相比,CAP在保持分词器(及其重建质量)固定的情况下,持续提高了任务成功率。我们的分析进一步表明,这些增益并不能由更高的分词准确率或策略头的单独变化来解释。相反,重用分词器码本为策略提供了关于分词器跨分词学习到的潜在结构的有价值信息,使得分词预测错误在动作空间中更为良性,并改善了策略主干学习到的表示。这些结果表明,动作分词器学习了超越离散目标的有用的动作感知潜在结构,这些结构应在训练下游策略时予以保留。
英文摘要
Action tokenization converts continuous robot actions into discrete symbols that can be modeled autoregressively. However, existing tokenizer-based policies typically ignore the tokenizer's learned latent code structure: after tokenization, the policy treats tokens as unrelated class indices and learns a new classifier from scratch. We show that this discarded structure is valuable. We introduce Codebook-Aligned Prediction (CAP), a method that directly reuses the tokenizer's code vectors as policy class prototypes while leaving the tokenizer and policy backbone otherwise unchanged. Across four quantizer families, three simulation benchmarks, and two real-robot tasks, CAP consistently improves task success over standard token classification heads while holding the tokenizer (and therefore its reconstruction quality) fixed. Our analysis further shows that these gains are not explained by higher token accuracy or changes in the policy head alone. Instead, reusing the tokenizer codebook provides the policy with valuable information about the tokenizer's learned latent structure across tokens, making token prediction errors more benign in action space and improving the representations learned by the policy backbone. These results suggest that action tokenizers learn useful action-aware latent structure beyond discrete targets that should be preserved when training downstream policies.