arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28865cs.CVcs.RO

动作表示中的方向-尺度分解:重新思考视觉-语言-动作模型的标记化对象

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

Yufei Duan, Hang Yin, Alberta Longhini, Chao Tang, Danica Kragic

首次发表
浏览论文内容

中文总结 AI 辅助

针对离散标记VLA模型,提出方向-尺度分解(DSD)动作表示,将平移和旋转增量分解为方向与尺度,在LIBERO、SimplerEnv及真实机器人上提升成功率,缓解混合数据集训练性能下降。

中文摘要 AI 辅助

动作表示在离散标记视觉-语言-动作(VLA)学习中扮演核心角色,但尚未得到充分研究。在传统的位姿增量表示下,动作标记对执行速度和数据集特定的归一化敏感,可能掩盖跨演示和数据集共享的几何结构。我们引入方向-尺度分解(DSD),一种在标记化之前将平移和旋转增量分解为方向和尺度分量的动作表示。DSD隔离运动方向,同时将幅度保留在单独的尺度通道中。我们在模拟和真实世界操作中,在单数据集和混合数据集训练下,使用均匀分箱(BIN)和BEAST(一种基于B样条的标记器)评估DSD。在LIBERO上,DSD使用两种标记器均提高了平均成功率。在SimplerEnv上,在混合数据集训练下,DSD-BIN在总体成功率上比BIN高出10.3个百分点。真实机器人实验进一步显示了在有和没有机器人预训练的情况下的增益。这些结果支持DSD作为离散标记VLA模型的有效动作表示,并表明其在大型多样化数据集混合训练时缓解性能下降的潜力。我们的项目页面及额外资源可在以下网址获取:此https URL

英文摘要

Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normalization, potentially obscuring geometric structure shared across demonstrations and datasets. We introduce Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization. DSD isolates motion direction while retaining magnitudes in separate scale channels. We evaluate DSD with uniform binning (BIN) and BEAST, a B-spline-based tokenizer, in simulation and real-world manipulation under both single-dataset and mixed-dataset training. On LIBERO, DSD improves average success rates with both tokenizers. On SimplerEnv, DSD-BIN outperforms BIN by 10.3 percentage points in overall success rate under mixed-dataset training. Real-robot experiments further show gains both with and without robotics pretraining. These results support DSD as an effective action representation for discrete-token VLA models and suggest its potential to mitigate performance degradation when training on large and diverse dataset mixtures. Our project page with additional resources is available at https://vla-dsd.github.io/

发表机构

  • University of Copenhagen(哥本哈根大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑