arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07361cs.ROcs.CV

驾驶视觉-语言-动作模型中规划 token 的深度探测与剪枝

Depth-Wise Probing and Pruning of the Planning Token in a Driving Vision-Language-Action Model

Harisankar Babu, Benjamin Coors, Christopher Lang, Hendrik Berkemeyer, Tamim Asfour, Simon Foell

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对驾驶 VLA 模型的规划 token,通过探测其 32 个解码器层的信号并剪枝部分层,在误差小幅增加下实现 1.33 倍解码器加速,验证了规划信息早期存在但格式适配问题。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型将驾驶决策通过深度语言模型传递,但动作本身需要多少层深度尚不明确。本文研究了一个代表性驾驶 VLA,其全部规划由单个规划 token 承载,生成式规划器会将该 token 解码为轨迹。借用该规划器作为轨迹空间 logit 透镜,我们从 32 个解码器层中的每一层解码规划 token,并测量两个信号:导航指令的线性可解码性,以及与冻结原生规划器的轨迹兼容性。诊断结果显示,语义意图可在早期线性解码:指令探测准确率在第一个解码器层后达到 97.7%,而随机概率为 16.7%;相比之下,与冻结原生规划器的兼容性随深度逐渐提升,开环平均 L2 距离仅在最终层达到最小值 2.11 米。从第一层学习的读出模块可弥补大部分差距,表明规划信息已在早期存在,但尚未以部署规划器所需的格式表示。按规划 token 引发的角度偏差对解码器层排序,可在相对开环误差增加约 5% 的情况下移除 32 个层中的 8 个,实现了测得的 1.33 倍解码器加速。在评估的样本量下,未统计到特定类别性能下降。这些发现仅适用于评估的 ORION 检查点和 Bench2Drive 设置。

英文摘要

Vision-language-action (VLA) models route driving decisions through a deep language model, but it is unclear how much of that depth the action itself requires. We study a representative driving VLA whose entire plan is carried by a single planning token that a generative planner decodes into a trajectory. Borrowing the planner as a trajectory-space logit lens, we decode the planning token from every one of the 32 decoder layers and measure two signals: the linear decodability of the navigation command and trajectory compatibility with the frozen native planner. Our diagnostic shows that semantic intent is linearly decodable early: command-probe accuracy reaches 97.7\% after the first decoder layer, compared with 16.7\% chance. In contrast, compatibility with the frozen native planner improves gradually across depth, with open-loop Avg-L2 reaching its minimum of 2.11\,m only at the final layer. Learned readouts from the first layer recover much of this gap, indicating that planning information is already present early but is not yet represented in the format expected by the deployed planner. Ranking decoder layers by the angular deviation they induce in the planning token permits removal of 8 of 32 layers within an approximately 5\% relative open-loop error increase and yields a measured 1.33$\times$ decoder speedup. At the evaluated sample size, no family-specific degradation is statistically resolved. These findings are limited to the evaluated ORION checkpoint and Bench2Drive setup.

发表机构

  • Robert Bosch GmbH(罗伯特·博世有限公司)
  • Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑