超越数据缩放:面向视觉-语言-动作模型的以表示为中心的持续预训练
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
浏览论文内容
中文总结 AI 辅助
针对VLA模型数据扩展的瓶颈,提出以表示为中心的持续预训练方法VLAct,在多具身机器人数据上训练后,在多项基准任务中超越现有工业级系统,在未见具身迁移任务中仅用20%数据即优于全数据基线,为VLA进展提供新方向。
中文摘要 AI 辅助
扩展机器人数据对于构建通用视觉-语言-动作(VLA)模型至关重要,但机器人轨迹的扩展难度远高于网络规模的图像-文本数据,因为具身数据收集成本高昂且对物理世界的覆盖稀疏。这使得表示质量成为核心瓶颈:在固定的机器人数据预算下,持续预训练必须将有限的轨迹转化为可迁移的视觉-动作知识,而非仅拟合动作。我们提出VLAct,这是一种面向VLA的VLM骨干模型,在针对特定任务进行微调前,已在广泛、异构、多具身的机器人数据上完成训练。VLAct通过VLM先验保留、多头连续动作协同监督以及部分统一的跨具身动作布局,在微调时允许使用特定任务的动作头,从而保留了广泛的VLM先验并鼓励不同具身之间共享动作语义。在模拟环境、真实世界以及未见具身的迁移任务中,VLAct在固定微调协议下始终提升下游性能:在LIBERO-Plus和RoboTwin 2.0上,VLAct超越了ABot-M0和LingBot-VLA等工业级VLA系统,分别达到82.6%和92.5%的成功率;在RoboDojo上,VLAct的成功率在所有策略中排名第六,且在两项指标上均优于所有明确指定的世界-动作模型(WAM)条目;最值得注意的是,在未见人形具身RoboCasa-GR1上,仅使用20%下游轨迹的VLAct,其性能优于使用全部数据的GR00T-N1.6基线。这些结果基于完全开源的数据和仅16个GPU的训练设置获得,表明以表示为中心的持续预训练可在适度的计算预算下提供极具竞争力的性能,是超越数据缩放的VLA进展的重要独立维度。
英文摘要
Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.