arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

驾驶VLA模型到Class 8卡车的低数据效率适配

Data-Efficient Adaptation of a Driving VLA to Class 8 Trucks

Satyajeet Das, Aaron Buxbaum, Niels Joubert, Gaurav S. Sukhatme

arXiv 2609.38570首次发表:更新:

发表机构

University of Southern California; Stack AV(南加州大学; Stack AV公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出“先适配后转向”策略,通过微调VLA模型并引入流速度转向模块,在少量数据下高效适配Class 8卡车驾驶,显著降低轨迹误差。

AI 中文摘要

Class 8卡车在几何形状、动力学特性和机动要求上与乘用车存在差异。因此,为乘用车训练的视觉-语言-动作(VLA)模型不能直接迁移到Class 8卡车,尤其是在事故现场和施工区域等非结构化场景中。我们提出了一种“先适配后转向”策略,而非从头训练卡车驾驶VLA模型,该策略适配现成的VLA模型以在这些具有挑战性的场景中为Class 8卡车生成轨迹。在适配阶段,我们使用NVIDIA的Alpamayo 1.5作为基础模型,仅在其动作生成栈上针对数百个真实世界的施工和事故相关高速公路场景进行微调。在转向阶段,我们引入了流速度转向(FVS)模块,在保持适配后的VLA模型固定的同时进一步细化模型的预测。FVS是一个紧凑的、基于流时间条件的残差模块,它为用于在每一步生成时更新动作序列的动作空间流速度添加学习到的修正。在场景不相交的留出集上的开环评估中,与基础模型相比,针对性微调在6.4秒的整个时间范围内将单候选平均位移误差(ADE)和最终位移误差(FDE)减少了一半以上。使用相同的针对性演示,FVS进一步将微调模型的全时域ADE和FDE分别降低了13.9%和16.5%。在匹配的数据预算下,针对性监督比通用卡车驾驶监督产生的全时域ADE低19-26%,而针对性模型仍然与在约65倍多的通用卡车驾驶场景上微调的模型保持竞争力。这些结果支持“先适配后转向”策略用于Class 8卡车的低数据效率车辆域迁移。我们的项目网站可在https URL访问。

英文摘要

Class 8 trucks differ from passenger cars in geometry, dynamics, and maneuvering requirements. As a result, vision-language-action (VLA) models trained for passenger vehicles do not readily transfer to Class 8 trucks, particularly in unstructured scenarios such as accident scenes and construction zones. Rather than training a truck-driving VLA from scratch, we propose an adapt-then-steer strategy that adapts an off-the-shelf VLA to generate trajectories for Class-8 trucks in these challenging scenarios. In the adapt stage, we use NVIDIA's Alpamayo 1.5 as the base model, fine-tuning only its action-generation stack on a few hundred real-world construction and accident-related highway scenarios. In the steer stage, we introduce Flow Velocity Steering (FVS) to further refine the model's predictions while holding the adapted VLA fixed. FVS is a compact, flow-time-conditioned residual module that adds learned corrections to the action-space flow velocity used to update the action sequence at each generation step. In open-loop evaluation on a scenario-disjoint held-out set, targeted fine-tuning more than halves single-candidate average displacement error (ADE) and final displacement error (FDE) over the entire 6.4 s horizon compared to the base model. Using the same targeted demonstrations, FVS further reduces the fine-tuned model's full-horizon ADE and FDE by 13.9% and 16.5%, respectively. At matched data budgets, targeted supervision yields 19-26% lower full-horizon ADE than general truck-driving supervision, while the targeted model remains competitive with a model fine-tuned on approximately 65 times as many general truck-driving scenarios. These results support adapt-then-steer for data-efficient vehicle-domain transfer to Class 8 trucks. Our project website is available at https://truckvla.github.io.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑