arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Vela:基于自适应动作曲线参数化的视觉-语言-动作模型扩展

Vela: Scaling Vision-Language-Action Models with Adaptive Action Curve Parametrization

Yifan Li, Jiaxu Wang, Dongming Wu, Yicheng Jiang, Ryan Ji, Xiangyu Yue, Yanwei Fu

arXiv 2610.05230首次发表:更新:

发表机构

Fudan University; Shanghai Innovation Institute; The Chinese University of Hong Kong; NovaXBot(复旦大学; 上海创新研究院; 香港中文大学; NovaXBot)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Vela提出基于样条的连续动作表示,通过自适应时间分辨率解决固定输出预算的局限,在仿真和真实任务中验证了其作为具身基础模型基础的潜力。

AI 中文摘要

大多数视觉-语言-动作模型将未来运动表示为固定速率的动作块,将时间分辨率和预测范围绑定到固定的输出预算。这种逐点表示在高度相关的相邻动作上浪费容量,将时间连续性和平滑性留给隐式学习,并迫使在长时程覆盖与接触丰富操作所需的局部精度之间进行权衡。为解决这些局限性,我们引入了Vela,一个将未来机器人行为表示为连续轨迹的视觉-语言-动作基础模型。Vela结合了基于紧凑样条的动作表示、运动相关的时间支持以及用于异构本体的共享动作接口,使得固定的输出预算能够跨运动调整其时间分辨率。我们在大规模多本体机器人数据上预训练Vela,并在LIBERO-X、EBench以及两个真实世界长时程任务(蛋饼烹饪和土豆切丝)上进行了评估,在仿真和物理操作中取得了有前景的结果。这些结果凸显了连续动作表示作为未来具身基础模型基础的潜力。项目页面和更多结果:此HTTPS URL。

英文摘要

Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves temporal continuity and smoothness to be learned implicitly, and forces a tradeoff between long-horizon coverage and the local precision required for contact-rich manipulation. To address these limitations, we introduce Vela, a vision-language-action foundation model that represents future robot behavior as continuous trajectories. Vela combines a compact spline-based action representation with motion-dependent temporal support and a shared action interface for heterogeneous embodiments, allowing a fixed output budget to adapt its temporal resolution across motions. We pretrain Vela on large-scale multi-embodiment robot data and evaluate it on LIBERO-X, EBench, and two real-world long-horizon tasks, egg-cake cooking and potato shredding, obtaining promising results across simulation and physical manipulation. These results highlight the potential of continuous action representations as a foundation for future embodied foundation models. Project page and more results: https://clementine24.github.io/Vela/ .

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑