arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09016cs.AI

PAIR:弥合视觉-语言-动作模型中的感知与行动

PAIR: Bridging Perception and Action in Vision-Language-Action Models

Kaixi Feng, Guoheng Sun, Ang li

首次发表
浏览论文内容

中文总结 AI 辅助

PAIR框架通过掩码动作自编码器和桥接模块学习共享的感知-动作表征,提升连续动作VLA模型在LIBERO、CALVIN及真实任务中的成功率。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型将视觉观察和语言指令映射为连续的机器人动作。这一任务要求从描述场景和指令的表征过渡到支持动作生成的表征。许多连续动作VLA模型将这种过渡隐式化,并主要通过最终的动作预测损失对其进行监督。我们提出了PAIR,一个在这两个空间之间学习共享的感知-动作表征的框架。在训练期间,掩码动作自编码器将专家动作片段编码为与时间范围对齐的动作潜在令牌。桥接模块从当前的视觉-语言表征中提取任务相关特征。PAIR将这些特征与动作潜在令牌对齐,形成保留任务信息并捕获专家动作结构的桥接令牌。随后,桥接令牌被投影到动作令牌空间,并注入初始动作令牌,为动作专家细化提供一个即时的动作起始点。在推理时,自编码器被移除,桥接令牌仅由当前观察和指令生成。在LIBERO、LIBERO-Plus和CALVIN ABC-D上的实验表明,所评估的OpenVLA-OFT和VLA-Adapter模型均获得了性能提升。在LIBERO-Plus上,PAIR将VLA-Adapter的成功率从59.1%提升至64.2%。在CALVIN上,PAIR将VLA-Adapter的平均完成序列长度从4.42提升至4.53。在七个真实世界任务中,PAIR将OpenVLA-OFT的成功率从51.4%提升至65.0%。表征分析表明,桥接令牌保留了任务信息,同时在动作专家细化之前使连续动作信息变得可访问。这些结果支持共享的中间表征作为连续动作VLA中感知与行动之间的有效接口。

英文摘要

Vision-language-action (VLA) models map visual observations and language instructions to continuous robot actions. This task requires a transition from representations that describe the scene and instruction to representations that support action generation. Many continuous-action VLAs leave this transition implicit and supervise it mainly through the final action-prediction loss. We introduce PAIR, a framework that learns a shared perception-action representation between these two spaces. During training, a Masked Action Autoencoder encodes expert action chunks into horizon-aligned Action Latent Tokens. A Bridge Module extracts task-relevant features from the current visual-language representations. PAIR aligns these features with the Action Latent Tokens to form Bridge Tokens that preserve task information and capture the structure of expert actions. The Bridge Tokens are then projected into the action-token space and injected into the initial Action Tokens, providing an action-ready starting point for Action Expert refinement. At inference, the autoencoder is removed, and the Bridge Tokens are generated only from the current observation and instruction. Experiments on LIBERO, LIBERO-Plus, and CALVIN ABC-D show gains for the evaluated OpenVLA-OFT and VLA-Adapter models. On LIBERO-Plus, PAIR raises VLA-Adapter's success rate from 59.1% to 64.2%. On CALVIN, it increases VLA-Adapter's average completed sequence length from 4.42 to 4.53. Across seven real-world tasks, PAIR raises OpenVLA-OFT's success rate from 51.4% to 65.0%. Representation analyses show that Bridge Tokens retain task information while making continuous-action information accessible before Action Expert refinement. These results support a shared intermediate representation as a useful interface between perception and action in continuous-action VLAs.

发表机构

  • University of Maryland, College Park(马里兰大学学院公园分校)

机构由 AI 辅助整理,请以论文原文为准。

↑