发表机构
Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; AI Lab, The Yangtze River Delta; Li Auto Inc.; School of Artificial Intelligence, Beijing University of Posts and Telecommunications; University College London; University of Edinburgh(中国科学院自动化研究所; 中国科学院大学; 长三角人工智能实验室; 理想汽车; 北京邮电大学人工智能学院; 伦敦大学学院; 爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控实验揭示VLA动作头性能主要由初始化决定,提出对齐的动作头与容量扩展策略,并构建了高效紧凑模型EffVLA,在低延迟下匹配或超越开源VLA。
AI 中文摘要
视觉-语言-动作(VLA)模型结合了预训练的视觉编码器、语言骨干网络和动作头,但在受控且延迟配对的条件下,它们各自的相对贡献尚未被确立。我们固定骨干网络家族(SigLIP2和Qwen2.5)和训练流程,扫描动作头设计和模块规模,并将每种配置与实测的设备端延迟配对。该研究得出三项发现。首先,动作头的性能主要由初始化而非解码器架构、损失函数或推理预算决定:将语言骨干网络的最后几层Transformer复制到动作头中是最大的单一杠杆,且不增加延迟成本,并且是唯一在每个模块规模下都有帮助的轴。对齐也解释了其他轴:流匹配和更重的解码器仅在动作头未对齐时有效,一旦对齐则效果反转,额外的推理次数没有可测量的收益;表达能力似乎替代了缺失的对齐。我们将此解读为表示迁移:对齐的动作头持续关注指令中的物体名词,并在权重空间中保持与骨干网络接近,而非从头重新学习动作。由于我们仅通过初始化达到对齐,我们将其作为最能组织测量结果的解释,而非已证明的因果,并指出能解决此问题的对照实验。其次,容量仅在对齐后才有效:对齐的动作头是扩展回报最高的模块。第三,这些回报在接近当今π系列VLA已使用的规模附近急剧下降,因此进一步增长对域内精度提升甚微,却增加延迟。这些成果定义了EffVLA,一个紧凑模型,在标准LIBERO上匹配最强的开源VLA,在大多数LIBERO-Plus扰动轴上以更低延迟领先,并以不变的配方迁移到真实的SO-ARM101机械臂上。
英文摘要
Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has not been established under controlled, latency-paired conditions. We fix the backbone families (SigLIP2 and Qwen2.5) and the training pipeline, sweep action-head design and module scale, and pair each configuration with measured on-device latency. The study yields three findings. First, action-head performance is governed primarily by initialization rather than decoder architecture, loss, or inference budget: copying the last transformer layers of the language backbone into the head is the single largest lever, at no latency cost, and the only axis that helps at every module scale. Alignment also explains the other axes: flow matching and a heavier decoder pay off only while the head is misaligned and reverse once it is aligned, and extra inference passes give no measurable benefit; expressiveness appears to substitute for missing alignment. We read this as representation transfer: the aligned head keeps attending to the instruction's object nouns and stays close to the backbone in weight space rather than relearning to act from scratch. Because we reach alignment only through initialization, we offer this as the account that best organizes the measurements, not a demonstrated cause, and name the control that would settle it. Second, capacity pays only after alignment: the aligned action head is the highest-return module to scale. Third, those returns diminish sharply near the size today's $π$-series VLAs already use, so further growth buys little in-domain accuracy for its latency. These specify EffVLA, a compact model matching the strongest open-source VLAs on standard LIBERO, leading on most LIBERO-Plus perturbation axes at lower latency, and transferring to a real SO-ARM101 arm with the recipe unchanged.