arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlashDrive:面向自动驾驶的Flash视觉-语言-动作推理

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

Zekai Li, Yihao Liang, Hongfei Zhang, Jian Chen, Yesheng Liang, Zhijian Liu

arXiv 2608.12932首次发表:更新:

发表机构

Princeton; UC San Diego(普林斯顿大学; 加州大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FlashDrive通过算法-系统协同设计解决VLA推理的四个级联瓶颈,使Alpamayo 1.5-10B的端到端自动驾驶延迟降4.7倍,推理频率提升至6.6Hz,逼近实时部署。

AI 中文摘要

视觉-语言-动作(VLA)模型有望为自动驾驶带来端到端推理能力,但其计算成本过高,无法满足实时控制需求。核心挑战具有结构性:VLA推理并非单一瓶颈,而是由四个瓶颈级联构成:视觉编码在重叠视频帧上浪费算力;语言模型预填充会重复计算可从上一时间步继承的上下文;推理令牌虽熵值低却仍串行生成;流匹配去噪对非均匀速度场采用均匀算力。单独解决任一阶段均无法触及其他阶段。我们提出FlashDrive,一种算法-系统协同设计框架,同时针对全部四个阶段。核心见解是每个瓶颈均可采用不同的轻量算法捷径:时间重叠支持跨帧的流式KV缓存复用;驾驶领域推理的低单令牌熵与强块内相关性,使非自回归扩散草稿器对投机解码极为有效;速度场结构(端点处尖锐、中间平坦)允许自适应步缓存,将算力集中于关键区域。结合系统级CUDA Graph编译与内核融合,这些技术效果叠加。在Alpamayo 1.5-10B(采用W4A8量化)上应用FlashDrive后,端到端延迟从717ms降至151ms(降低4.7倍),同时精度基本未变:minADE6@6.4s仅偏移0.08m,minADE1有所提升,仿真中的闭环碰撞率与 off-road 率均改善。通过在单GPU上将100亿参数推理VLA的推理频率从1.4Hz提升至6.6Hz,FlashDrive大幅推动端到端自动驾驶向实时部署迈进。

英文摘要

Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4~Hz to 6.6~Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.

Comments15 pages; 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑