发表机构
Princeton; UC San Diego(普林斯顿大学; 加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FlashDrive通过算法-系统协同设计解决VLA推理的四个级联瓶颈,使Alpamayo 1.5-10B的端到端自动驾驶延迟降4.7倍,推理频率提升至6.6Hz,逼近实时部署。
AI 中文摘要
视觉-语言-动作(VLA)模型有望为自动驾驶带来端到端推理能力,但其计算成本过高,无法满足实时控制需求。核心挑战具有结构性:VLA推理并非单一瓶颈,而是由四个瓶颈级联构成:视觉编码在重叠视频帧上浪费算力;语言模型预填充会重复计算可从上一时间步继承的上下文;推理令牌虽熵值低却仍串行生成;流匹配去噪对非均匀速度场采用均匀算力。单独解决任一阶段均无法触及其他阶段。我们提出FlashDrive,一种算法-系统协同设计框架,同时针对全部四个阶段。核心见解是每个瓶颈均可采用不同的轻量算法捷径:时间重叠支持跨帧的流式KV缓存复用;驾驶领域推理的低单令牌熵与强块内相关性,使非自回归扩散草稿器对投机解码极为有效;速度场结构(端点处尖锐、中间平坦)允许自适应步缓存,将算力集中于关键区域。结合系统级CUDA Graph编译与内核融合,这些技术效果叠加。在Alpamayo 1.5-10B(采用W4A8量化)上应用FlashDrive后,端到端延迟从717ms降至151ms(降低4.7倍),同时精度基本未变:minADE6@6.4s仅偏移0.08m,minADE1有所提升,仿真中的闭环碰撞率与 off-road 率均改善。通过在单GPU上将100亿参数推理VLA的推理频率从1.4Hz提升至6.6Hz,FlashDrive大幅推动端到端自动驾驶向实时部署迈进。
英文摘要
Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4~Hz to 6.6~Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.
Comments15 pages; 8 figures