arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27513cs.RO

行为对齐的动作标记化用于机器人策略学习

Behavior-Aligned Action Tokenization for Robot Policy Learning

Junbo Dong, Ze Chen, Zhendong Xie, Junjie Li, Lixin Xu, Xuemin Chi, Yiming Song, Zhaoyuan Ma

首次发表
浏览论文内容

中文总结 AI 辅助

针对机器人策略学习中动作标记缺乏行为对应监督的问题,提出BAAT方法,利用软动态时间规整对齐动作块并联合量化重建,使相似运动共享表示,在仿真和真实任务上显著提升策略成功率。

中文摘要 AI 辅助

自回归机器人策略通过从观测中预测离散动作标记来学习连续控制。不同任务通常共享局部运动,然而现有标记器中对演示之间的行为对应关系提供的显式监督有限。因此,具有不同时机的运动尽管遵循相似模式,却可能缺乏共享表示。我们提出了行为对齐的动作标记化(BAAT),该方法使用软动态时间规整(Soft-DTW)来选择对应的动作块,并联合重建对齐其量化坐标。该目标鼓励不同任务中的相似运动占据邻近的量化表示,同时保留可执行的动作细节。一个基于历史条件的扩散解码器从这些标记中重建连续动作块,下游的自回归策略学习预测这些标记。我们在三个仿真基准的选定任务和两个真实机器人任务上评估了BAAT。BAAT的平均仿真成功率达到约45.2%,比OAT高出约7.2个百分点。在受控的LIBERO-All对齐消融实验中,策略成功率从70.2%上升到79.0%,而轨迹重放成功率下降。这些结果支持将行为对应关系作为监督信号,用于组织动作标记器中的共享运动结构,并改进下游机器人策略学习。

英文摘要

Autoregressive robot policies learn continuous control by predicting discrete action tokens from observations. Different tasks often share local motions, yet behavioral correspondence across demonstrations receives limited explicit supervision in existing tokenizers. Motions with different timing can therefore lack a shared representation despite following similar patterns. We propose Behavior-Aligned Action Tokenization (BAAT), which uses soft dynamic time warping (Soft-DTW) to select corresponding action chunks and aligns their quantized coordinates jointly with reconstruction. This objective encourages similar motions across tasks to occupy nearby quantized representations while retaining executable action detail. A history-conditioned diffusion decoder reconstructs continuous action chunks from these tokens, and a downstream autoregressive policy learns to predict them. We evaluate BAAT on selected tasks from three simulation benchmarks and two real robot tasks. BAAT achieves a mean simulation success rate of approximately 45.2%, exceeding OAT by approximately 7.2 percentage points. In the controlled LIBERO-All alignment ablation, policy success rises from 70.2% to 79.0% while trajectory replay success decreases. These results support behavioral correspondence as supervision for organizing shared motion structure in action tokenizers and improving downstream robot policy learning.

发表机构

  • Southern University of Science and Technology(南方科技大学)
  • Dexmal
  • National University of Singapore(新加坡国立大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

↑