arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

离散思维到连续动作:面向端到端自动驾驶的隐式对齐规划

Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving

Ruoyu Yao, Yusen Xie, Qingzhao Liu, Pei Liu, Zewei Yang, Yipeng Zhu, Xiaolong Wang, Jun Ma

arXiv 2609.04070首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Central Media Technology Institute, Huawei(香港科技大学(广州); 华为中央媒体技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出具备隐式对齐规划的VLA框架LaPla,通过残差VQ-VAE消除量化误差,在nuScenes基准和AlpaSim模拟器上分别实现长时程L2误差降低、驾驶成功率提升的效果。

AI 中文摘要

弥合视觉语言模型的离散推理与自动驾驶受物理约束的连续特性之间的差距仍是一项重大挑战。本研究提出LaPla,一种具备隐式对齐规划的统一视觉-语言-动作(VLA)框架,用于将语义理解无缝落地到精准的运动执行中。我们首先设计了一种基于残差向量量化变分自编码器(VQ-VAE)的动作分词器,捕捉车辆运动学并将轨迹特征编码到结构化隐式空间。LaPla未采用必然引入量化误差的离散码本查找,而是将该表示作为物理先验,弥合高维语义与原始动作空间之间的模态差距。具体而言,给定整合多视角图像、历史动作及文本指令的多模态输入,LaPla纳入并发动作查询,在单次前向传播中因果关注多模态上下文,将隐藏状态直接投影到预训练VQ-VAE隐式空间。冻结的解码器随后将这些连续隐式转化为动作,有效消除量化误差,确保轨迹符合物理规律,同时规避耗时的自回归生成。在nuScenes基准上的大量实验表明,LaPla具备具竞争力的开环性能,与最先进的VLA方法相比,长时程L2误差降低15.52%。在NVIDIA AlpaSim模拟器上的闭环评估进一步证实其保障驾驶平顺性的卓越能力,成功率提升33.34个百分点,且推理延迟显著降低。

英文摘要

Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrained nature of autonomous driving remains a significant challenge. In this work, we introduce LaPla, a unified Vision-Language-Action (VLA) framework featuring latent-aligned planning to seamlessly ground semantic understanding in precise motion execution. We first design an action tokenizer based on a residual vector-quantized variational autoencoder (VQ-VAE), capturing vehicle kinematics and encoding trajectory features into a structured latent space. Rather than discrete codebook lookups that inevitably introduce quantization errors, LaPla repurposes this representation as a physical prior to bridge the modality gap between high-dimensional semantics and the raw action space. Specifically, given multimodal inputs integrating multi-view images, historical actions, and textual instructions, LaPla incorporates concurrent action queries to causally attend to the multimodal context in a single forward pass, projecting hidden states directly into the pretrained VQ-VAE latent space. The frozen decoder then translates these continuous latents into actions, effectively eliminating quantization errors and ensuring physically plausible trajectories while bypassing time-consuming autoregressive generation. Extensive experiments on the nuScenes benchmark demonstrate that LaPla achieves competitive open-loop performance, reducing long-horizon L2 error by 15.52% compared to state-of-the-art VLA methods. Closed-loop evaluations on the NVIDIA AlpaSim simulator further confirm its superior capability in ensuring smooth driving progress, improving the success rate by 33.34 percentage points with significantly reduced inference latency.

Comments8 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑