arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ATI-VLA:通过可操作对齐然后自适应注入的以动作中心的预测视觉-语言-动作模型

ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection

Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong Yu

arXiv 2610.01741首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; Institute for Artificial Intelligence, Great Bay University; Shenzhen Loop Area Institute; Macao Polytechnic University; Shenzhen Technology University; Dongguan Key Laboratory for Intelligence and Information Technology(哈尔滨工业大学(深圳); 大湾区大学人工智能研究院; 深圳河套学院; 澳门理工大学; 深圳技术大学; 东莞市智能与信息技术重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对预测式VLA模型因模态错位和优化冲突而性能不佳的问题,提出ATI-VLA框架,通过共享码本对齐表示并自适应注入预测潜在变量,在仿真和真实任务上达到最先进性能且收敛更快。

AI 中文摘要

预测式视觉-语言-动作(VLA)模型旨在通过未来观测或世界动态预测来改进机器人操作。然而,现有方法往往未能实现这一潜力,其性能不如直接的动作预测模型。我们认为这些局限性源于观测与动作之间的模态错位,以及联合优化冲突使得学习偏离以动作中心的目标。为此,我们引入了ATI-VLA,一种通过可操作对齐然后自适应注入的以动作中心的预测视觉-语言-动作框架。具体而言,它采用两步设计:1)通过共享码本进行可操作表示对齐。它通过统一的码本将预测观测和动作表示映射到共享的离散潜在空间,从而使预测观测潜在变量可直接用于动作生成,并缓解模态错位。2)预测潜在变量的以动作中心的自适应注入。在此基础上,它通过轻量级自适应旁路将预测观测潜在变量作为显式预测先验注入动作解码,从而在单一以动作中心的目标下实现自适应预测引导。在仿真和真实世界机器人任务上的大量实验表明,ATI-VLA实现了最先进的性能,且收敛速度更快。

英文摘要

Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.

CommentsAccepted to NeurIPS 2026. Project page: https://jiutian-vl.github.io/ATI-VLA-page/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑