发表机构
École des Mines de Paris; Stellantis(巴黎矿业学院; 斯泰兰蒂斯集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对驾驶VLAs的跨embodiment零样本迁移问题,提出BEV-Forcing辅助目标,通过多数据集训练提升性能,发现辅助任务收益随训练embodiment数量增加而降低。
AI 中文摘要
视觉-语言-动作模型(VLAs)通过利用多模态预训练实现指令跟随、视觉推理和场景级泛化,在自动驾驶领域展现出强大潜力。在机器人操纵领域,跨多个机器人设置扩展VLA微调(尤其是统一不同 embodiment 的表征)已被证实可提升数据集内性能与跨 embodiment 泛化能力;但在自动驾驶领域,VLAs大多仍在单一数据集上训练,极少被评估对未见过的数据集和相机 rig 的零样本迁移能力,且盲目增加训练数据量未必能提升已见过 embodiment 的性能。为解决这些问题,本研究针对驾驶任务开展多数据集训练,并提出BEV-Forcing这一辅助目标,该目标将专用鸟瞰图(Bird's-Eye-View)模型的地面平面物体布局信息迁移至VLA骨干网络。通过促使模型通过共享BEV空间接口表征物体位置,研究表明,在少量相机rig上训练时,BEV-Forcing这类辅助任务可同时提升分布内与分布外性能;然而,随着训练中embodiment数量增加,该辅助任务的收益会降低,这一现象为相关文献中“仅扩大训练多样性会导致新方法收益下降”的观点提供了证据,也促使研究需结合数据规模扩展来呈现结果。
英文摘要
Vision-Language-Action models (VLAs) have shown strong potential in autonomous driving by leveraging multimodal pretraining for instruction following, visual reasoning, and scene-level generalization. In robotic manipulation, scaling VLA fine-tuning across multiple robot setups--especially when unifying representations across embodiments--has been shown to improve in-dataset performance and cross-embodiment generalization; in autonomous driving, however, VLAs remain largely trained on individual datasets and are rarely evaluated for zero-shot transfer to unseen datasets and camera rigs; furthermore naively adding more datasets to the training data does not necessarily lead to better performance within seen embodiments. To address these problems, we study multi-dataset training for the driving task and BEV-Forcing, an auxiliary objective that transfers ground-plane object-layout information from a specialized Bird's-Eye-View model into the VLA backbone. By encouraging the model to represent object position through a shared BEV spatial interface, we show that an auxiliary task such as BEV-Forcing can improve both in-distribution and out-of-distribution performance when training on a small number of camera rigs. As the number of training embodiments increases, however, the benefits of the auxiliary task are reduced; we present this as evidence that new techniques in the literature may see their benefits diminish when simply scaling up training diversity, which motivates presenting results taking into account data scaling.