骨干网络中的规划:DiffAdapterVLA用于驾驶视觉语言模型的原生连续轨迹生成
Planning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs
浏览论文内容
中文总结 AI 辅助
DiffAdapterVLA通过将轨迹令牌注入VLM骨干后期层,实现与驾驶条件共同演化的连续轨迹规划,仅用轻量模块即达高效闭环规划。
中文摘要 AI 辅助
预训练的驾驶视觉语言模型(VLMs)将视觉、路线、语言和驾驶上下文整合为丰富的驾驶先验,但其表示目标与连续驾驶规划相分离。现有方法通常在VLM形成最终条件后才开始轨迹生成,将深度条件的计算排除在轨迹状态的逐步形成之外。我们提出DiffAdapterVLA,实现了骨干网络中的规划:它将显式轨迹令牌注入选定的VLM后期层,使轨迹状态进入骨干网络的前向计算,并在不同深度与驾驶条件共同演化。轻量级的逐层DiffAdapter将此计算组织为递归轨迹细化,而非对称联合注意力保留了从条件流到轨迹规划的定向引导。通过将规划置于现有骨干网络计算中,而非依赖独立的轨迹规划器,DiffAdapterVLA仅适配轻量级轨迹模块,将现有驾驶先验转化为高效的连续规划能力。NAVSIM结果表明,它使用少量可训练参数实现了高质量闭环规划和低端到端延迟,并证明在VLM后期层计算中联合演化轨迹状态和深度驾驶条件可有效实现连续轨迹规划。
英文摘要
Pretrained driving vision-language models (VLMs) integrate visual, route, language, and driving context into rich driving priors, yet their representation objectives remain separated from continuous driving planning. Existing methods typically begin trajectory generation only after the VLM has formed a final condition, leaving depth-wise condition computation outside the stepwise formation of trajectory state. We introduce DiffAdapterVLA, which realizes Planning in the Backbone: it injects explicit trajectory tokens into selected VLM late layers, bringing trajectory state into backbone forward computation, where it co-evolves with driving conditions at different depths. Lightweight layer-wise DiffAdapters organize this computation into recursive trajectory refinement, while asymmetric joint attention preserves directed guidance from the condition stream to trajectory planning. By placing planning within existing backbone computation rather than relying on an independent trajectory planner, DiffAdapterVLA adapts only lightweight trajectory modules to turn existing driving priors into efficient continuous planning capability. NAVSIM results show that it achieves high-quality closed-loop planning with low end-to-end latency using few trainable parameters, and demonstrate that jointly evolving trajectory state and depth-wise driving conditions in VLM late-layer computation effectively realizes continuous trajectory planning.
发表机构
- Wuhan University(武汉大学)
- Dongfeng Research & Development Institute(东风汽车研发总院)
机构由 AI 辅助整理,请以论文原文为准。