arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考语言在自动驾驶高效视觉-语言-动作(VLA)模型中的作用:迈向更智能、可信的驾驶

Rethinking Language's Role in Efficient VLA for Autonomous Vehicles: Toward Smarter, Trustworthy Driving

Tongfei Guo, Lili Su

arXiv 2608.30144首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对自动驾驶VLA模型的高推理开销问题,提出语言残差分类法梳理高效语言使用方法,在多基准上分析适配性,将发布相关代码仓库。

AI 中文摘要

视觉-语言-动作(VLA)模型正通过语言统一感知、推理与控制,重塑自动驾驶(AD)领域,实现语义落地、可解释决策及更优长尾泛化。但车载端语言相关开销高昂:延迟与内存预算紧张,且自回归解码本质为顺序处理。本研究将核心问题重新定义为推理阶段语言应在何时、何处发挥作用,因推理开销在每帧部署时重复产生,而训练开销仅需支付一次。我们提出语言残差分类法,按推理阶段对语言的使用方式将方法分为四类:仅训练阶段监督(L1)、潜在非文本推理(L2)、条件调用(L3)及逐帧全生成(L4)。我们综述代表性方法并在五个部署维度(延迟、参数、内存、浮点运算量、令牌)上对其标记,在主要开环与闭环驾驶基准(如nuScenes、NAVSIM、Bench2Drive)上分析这些方法。我们进一步梳理NLP/大语言模型(LLM)中的高效方法如何适配自动驾驶,明确驱动这些适配的约束与动机。一个持续更新的代码仓库将在Github上发布。

英文摘要

Vision-Language-Action (VLA) models are reshaping autonomous driving (AD) by unifying perception, reasoning, and control through language, enabling semantic grounding, interpretable decisions, and better long-tail generalization. But language is expensive onboard: latency and memory budgets are tight, and autoregressive decoding is inherently sequential. This work reframes the central question as when and where language should act at inference, since inference cost recurs at every deployed frame while training cost is paid once. We introduce the Language Residue taxonomy to organize methods by their inference-time use of language: train-time-only supervision (L1), latent non-textual reasoning (L2), conditional invocation (L3), and full per-frame generation (L4). We review representative methods and tag each across five deployment axes (latency, parameters, memory, FLOPs, tokens), analyzing them on major open- and closed-loop driving benchmarks (e.g., nuScenes, NAVSIM, Bench2Drive). We further trace how efficient methods from NLP/LLM are adapted in AD, identifying the constraints and motivations driving these adaptations. A continuously updated repository will be available at Github.

CommentsAccepted to EMNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑