arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00237cs.CVcs.RO

隐式质心引导:用于指令对齐自动驾驶的单次无分类器引导

Latent-Centroid Steering: Single-Pass Classifier-Free Guidance for Command-Aligned Autonomous Driving

Meibo Hu, Jiamian Wang, Pichao Wang, Zhiqiang Tao

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉语言自动驾驶模型的指令跟随差距,提出单次引导机制LCS,降低推理延迟约50%,在Bench2Drive和nuScenes基准上提升指令遵循与驾驶性能。

中文摘要 AI 辅助

视觉语言模型(VLMs)近期成为端到端自动驾驶的有前景范式,使智能体能将多模态输入与高级导航指令直接映射为可执行轨迹。但实际应用中,这些模型存在持续的指令跟随差距:预测轨迹对导航指令的敏感性弱,导致关键决策点出现错误行为。我们将此问题归为条件策略崩溃的一种形式,即基于多模态轨迹分布的回归训练,促使模型依赖主导视觉先验,同时边缘化语言条件信号。为解决该问题,我们为基于回归的视觉语言驾驶提出了无分类器引导(CFG)的原理性公式。我们证明,CFG可解释为通过对比条件与无条件预测,在动作空间中分离指令诱导的残差,从而在推理时显式放大导航指令的效果。然而,标准的两次通过CFG会引入实时控制无法接受的延迟,并产生噪声的实例级引导方向。基于CFG的均值漂移解释,我们提出隐式质心引导(LCS),这是一种单次通过的引导机制,用类别级隐式偏移替代实例级残差。通过将条件表示投影到预计算的特定指令质心,LCS基于聚类几何执行类别级隐式引导,既更稳定又计算高效。我们证明,LCS将推理延迟降低约50%,同时在闭环(Bench2Drive)和开环(nuScenes)基准上实现更强的指令遵循和更优的驾驶性能。代码将公开。

英文摘要

Vision-language models (VLMs) have recently emerged as a promising paradigm for end-to-end autonomous driving, enabling agents to map multimodal inputs and high-level navigation instructions directly to executable trajectories. However, in practice, these models exhibit a persistent command-following gap: predicted trajectories often show weak sensitivity to navigation commands, resulting in incorrect behavior at critical decision points. We identify this issue as a form of conditional policy collapse, where regression-based training under multimodal trajectory distributions encourages the model to rely on dominant visual priors while marginalizing the language-conditioned signal. To address this issue, we introduce a principled formulation of classifier-free guidance (CFG) for regression-based vision-language driving. We show that CFG can be interpreted as isolating the instruction-induced residual in the action space by contrasting conditional and unconditional predictions, thereby explicitly amplifying the effect of the navigation command at inference time. However, a standard two-pass CFG introduces prohibitive latency for real-time control and produces noisy instance-level guidance directions. Building on a mean-shift interpretation of CFG, we propose Latent-Centroid Steering (LCS), a single-pass guidance mechanism that replaces instance-level residuals with class-level latent shifts. By projecting conditional representations toward precomputed command-specific centroids, LCS performs class-level latent steering based on cluster geometry that is both more stable and computationally efficient. We demonstrate that LCS reduces inference latency by approximately 50% while achieving stronger command adherence and improved driving performance on both closed-loop (Bench2Drive) and open-loop (nuScenes) benchmarks. Code are available at https://github.com/codingmlinprocess/LCS.

发表机构

  • Rochester Institute of Technology(罗切斯特理工学院)
  • NVIDIA(英伟达公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑