arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TOAST:面向自回归视觉-语言-动作模型的随机机器人动作分词方法

TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models

Keisuke Shirai, Tomohiro Motoda, Hanbit Oh, Ryoichi Nakajo, Roman Mykhailyshyn, Ryo Hanai, Shotaro Miwa, Yukiyasu Domae

arXiv 2610.00899首次发表:更新:

发表机构

AIST(产业技术综合研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出TOAST,一种随机动作分词方法,通过采样同一量化动作序列的多种分词结果来丰富离散监督,从而在有限数据下提升自回归机器人策略学习性能,实验显示成功率显著提升。

AI 中文摘要

自回归视觉-语言-动作模型通常将连续的机器人动作表示为离散的令牌序列,从而能够使用标准的下一令牌预测目标进行动作预测。FAST通过将包含不同时间频率的动作紧凑地编码为相对较少的令牌,显著改进了这种表示方式。然而,尽管这种压缩减少了自回归预测所需的动作令牌数量,但并不一定能提高从有限演示中学习策略的效率。特别是,FAST通常为每个量化动作序列分配一个确定性的分词结果,尽管多个令牌序列可以表示并解码为相同的机器人运动。我们研究了利用这种表示冗余是否能够改进策略学习。在本文中,我们提出了动作序列的随机分词方法(TOAST),这是一种随机动作分词方法,在策略训练过程中对同一量化动作序列的替代分词结果进行采样。这使离散监督信号多样化,同时保留底层机器人动作,且不需要额外的演示。在LIBERO上的实验表明,TOAST始终优于其确定性对应方法,且随着训练数据的减少,改进幅度增大,在仅使用1/16训练数据时,成功率提高了6.8个百分点。在四个真实机器人操作任务中,TOAST进一步将平均成功率比确定性对应方法提高了15.8个百分点。这些结果证明了随机动作分词对于自回归机器人策略学习的有效性,尤其是在训练数据有限的情况下。

英文摘要

Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few tokens. However, while such compression reduces the number of action tokens required for autoregressive prediction, it does not necessarily improve the efficiency of policy learning from limited demonstrations. In particular, FAST typically assigns a single deterministic tokenization to each quantized action sequence, although multiple token sequences can represent and decode to the same robot motion. We investigate whether exploiting this representational redundancy can improve policy learning. In this paper, we propose TOkenization of Action sequences with STochastic sampling (TOAST), a stochastic action tokenization method that samples alternative tokenizations of the same quantized action sequence during policy training. This diversifies the discrete supervision while preserving the underlying robot action and requires no additional demonstrations. Experiments on LIBERO show that TOAST consistently improves over its deterministic counterpart, with the improvement increasing as training data decreases, achieving a 6.8 point gain in success rate when only 1/16 of training data is available. Across four real-robot manipulation tasks, TOAST further improves mean success rate by 15.8 points over the deterministic counterpart. These results demonstrate the effectiveness of stochastic action tokenization for autoregressive robot policy learning, particularly when training data are limited.

CommentsProject page: https://kskshr.github.io/toast/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑