CLAP:通过语言-动作基础实现从视觉语言模型到视觉语言动作模型的直接适配
CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding
浏览论文内容
中文总结 AI 辅助
研究如何将预训练视觉语言模型直接转换为视觉语言动作模型,核心方法是提出CLAP,通过在数字动作序列前加自然语言描述来解决输出分布不匹配问题,单轮微调效果良好,还将发布多尺度紧凑VLA家族助力能力转移分析。
中文摘要 AI 辅助
视觉语言动作模型(VLA)从预训练的视觉语言模型(VLM)继承语义能力,但在机器人数据上的大规模训练后和架构修改会广泛重塑主干,难以分离VLM对控制的贡献。直接将预训练VLM转换为VLA且架构变化最小,能更清晰了解VLM能力如何跨模型规模转移。核心障碍是输出分布不匹配,预测动作作为纯数字令牌序列会使生成偏离VLM预训练语言分布。为此提出CLAP,在每个数字动作序列前加自然语言动作描述,在不修改主干架构的情况下因果调节精确动作令牌预测。单轮微调下,2B的CLAP在LIBERO上达到90.8%(比VLA - 0高14.9个百分点),并提高了在LIBERO - PRO上对语言、对象和空间扰动的鲁棒性。还将发布0.8B、2B和4B的CLAP作为开放权重、多尺度紧凑VLA家族,以实现对VLM到VLA能力转移的可控分析。
英文摘要
Vision-language-action models (VLAs) inherit semantic capabilities from pretrained VLMs, yet large-scale post-training on robot data and architectural modifications can reshape the backbone so extensively that it becomes difficult to isolate what the VLM contributes to control. Directly converting pretrained VLMs into VLAs with minimal architectural change offers a more transparent path to understanding how VLM capabilities transfer across model scales. The core obstacle is output-distribution mismatch: predicting actions as bare numeric token sequences moves generation away from the VLM's pretrained language distribution, degrading the capabilities we seek to preserve. To address this, we propose CLAP (Causal Language-Action Prediction), which prepends each numeric action sequence with a natural-language action description, causally conditioning precise action-token prediction on a language-action plan without modifying the backbone architecture. With single-epoch fine-tuning alone, 2B CLAP achieves 90.8% on LIBERO (+14.9 pt over VLA-0) and improves robustness on LIBERO-PRO under language, object, and spatial perturbations. We will release CLAP at 0.8B, 2B, and 4B as an open-weight, multi-scale compact VLA family from a single VLM lineage, enabling controlled analysis of VLM-to-VLA capability transfer.
发表机构
- Ochanomizu University(御茶水女子大学)
- University of Tokyo(东京大学)
- OMRON SINIC X Corp(欧姆龙(中国)有限公司)
机构由 AI 辅助整理,请以论文原文为准。