arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向驾驶视觉-语言-动作模型(VLAs)的跨 embodiment 零样本迁移研究

Towards Zero-Shot Transfer Across Embodiments For Driving VLAs

Caio Azevedo, Stefano Sabatini, Sascha Hornauer, Fabien Moutarde

arXiv 2609.02341首次发表:更新:

发表机构

École des Mines de Paris; Stellantis(巴黎矿业学院; 斯泰兰蒂斯集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对驾驶VLAs的跨embodiment零样本迁移问题,提出BEV-Forcing辅助目标,通过多数据集训练提升性能,发现辅助任务收益随训练embodiment数量增加而降低。

AI 中文摘要

视觉-语言-动作模型(VLAs)通过利用多模态预训练实现指令跟随、视觉推理和场景级泛化,在自动驾驶领域展现出强大潜力。在机器人操纵领域,跨多个机器人设置扩展VLA微调(尤其是统一不同 embodiment 的表征)已被证实可提升数据集内性能与跨 embodiment 泛化能力;但在自动驾驶领域,VLAs大多仍在单一数据集上训练,极少被评估对未见过的数据集和相机 rig 的零样本迁移能力,且盲目增加训练数据量未必能提升已见过 embodiment 的性能。为解决这些问题,本研究针对驾驶任务开展多数据集训练,并提出BEV-Forcing这一辅助目标,该目标将专用鸟瞰图(Bird's-Eye-View)模型的地面平面物体布局信息迁移至VLA骨干网络。通过促使模型通过共享BEV空间接口表征物体位置,研究表明,在少量相机rig上训练时,BEV-Forcing这类辅助任务可同时提升分布内与分布外性能;然而,随着训练中embodiment数量增加,该辅助任务的收益会降低,这一现象为相关文献中“仅扩大训练多样性会导致新方法收益下降”的观点提供了证据,也促使研究需结合数据规模扩展来呈现结果。

英文摘要

Vision-Language-Action models (VLAs) have shown strong potential in autonomous driving by leveraging multimodal pretraining for instruction following, visual reasoning, and scene-level generalization. In robotic manipulation, scaling VLA fine-tuning across multiple robot setups--especially when unifying representations across embodiments--has been shown to improve in-dataset performance and cross-embodiment generalization; in autonomous driving, however, VLAs remain largely trained on individual datasets and are rarely evaluated for zero-shot transfer to unseen datasets and camera rigs; furthermore naively adding more datasets to the training data does not necessarily lead to better performance within seen embodiments. To address these problems, we study multi-dataset training for the driving task and BEV-Forcing, an auxiliary objective that transfers ground-plane object-layout information from a specialized Bird's-Eye-View model into the VLA backbone. By encouraging the model to represent object position through a shared BEV spatial interface, we show that an auxiliary task such as BEV-Forcing can improve both in-distribution and out-of-distribution performance when training on a small number of camera rigs. As the number of training embodiments increases, however, the benefits of the auxiliary task are reduced; we present this as evidence that new techniques in the literature may see their benefits diminish when simply scaling up training diversity, which motivates presenting results taking into account data scaling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑