arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAC-VLA:用于视觉-语言-动作模型的上下文门控动作条件调节

CAC-VLA: Context-Gated Action Conditioning for Vision-Language-Action Models

Yifu Xiong, Wenhao Yu, Jiaxuan Lin, Bojun Zou, Jiahao Li, Lu Zhang, Yanyong Zhang, Jianmin Ji

arXiv 2607.04816首次发表:更新:

发表机构

University of Science and Technology of China (USTC); Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(中国科学技术大学; 合肥综合性国家科学中心人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉-语言-动作模型,提出上下文门控动作条件调节框架CAC-VLA,在视觉语言模型中学习轻量级潜在动作接口,训练模型预测潜在动作并通过上下文门调节动作专家,实验验证其有效性。

AI 中文摘要

视觉-语言-动作(VLA)模型是通用机器人操纵的有前景范式,但其表示未针对动作调节明确优化。我们提出CAC-VLA,一个上下文门控动作条件调节框架,在VLM内直接学习轻量级潜在动作接口,训练VLM预测潜在动作并通过上下文门调节动作专家,实验验证了其有效性。

英文摘要

Vision-Language-Action (VLA) models have become a promising paradigm for generalist robot manipulation, where visual-language representations are used to condition continuous action generation. However, these representations are not explicitly optimized for action conditioning, leaving the action expert to bridge the gap between multimodal understanding and precise motor control. Recent action-reasoning methods introduce additional modules to generate explicit action plans or action-space reasoning signals, demonstrating the benefit of action-level guidance but often requiring separate action-generation frameworks. We propose CAC-VLA, a Context-Gated Action Conditioning framework that learns a lightweight latent-action interface directly within the VLM. Instead of generating executable trajectories, CAC-VLA trains the VLM to predict coarse-to-fine latent actions, which are structured representations encoded from future action segments, and adaptively leverages them to condition the action expert via a context gate. This enables VLM-native action conditioning while calibrating the influence of latent-action guidance on expert action generation. Experiments on LIBERO, LIBERO-Plus, and CALVIN demonstrate the effectiveness of CAC-VLA, achieving 98.9% and 90.4% average success rates on LIBERO and LIBERO-Plus, respectively, and outperforming π0.5 by 9.3 percentage points on CALVIN.

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑