面向混合架构的视觉-语言-动作(VLA)高效块级并行推理
Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures
浏览论文内容
中文总结 AI 辅助
该研究针对VLA模型部署于车载平台的延迟与内存压力问题,提出混合CPU-GPU推理框架,通过块级划分实现资源调度,在Bench2Drive数据集及实车部署中均取得性能提升。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型正成为自动驾驶领域极具前景的范式,但将其部署在现有车载平台仍存在困难,因为它们会带来较高的推理延迟和显著的GPU端资源压力。在完整的自动驾驶栈中,这一问题更为突出:传统车载平台为模块化流水线配置,在多个规划相关功能被整合为统一的VLA模型后,原有的部分CPU预算被闲置,而视觉编码器和主推理路径仍将大部分计算与内存需求集中在GPU上。因此,在实际GPU内存约束下,直接将VLA与其他车载系统一同部署存在挑战。为解决该问题,我们提出一种适用于自动驾驶的混合CPU-GPU推理框架,具备灵活的资源调度能力。我们的设计在块级粒度对VLA主干进行划分,在GPU上执行视觉编码器和大语言模型(LLM)前缀,通过跨框架异步流水线将LLM后缀卸载到CPU,从而提供可调度边界,在异构处理器间重新分配计算与内存压力。我们在两个具有代表性的驾驶VLA模型Orion和MindDrive上评估了该框架:在Bench2Drive数据集上,我们的方法将Orion的平均延迟从521ms降至408.0ms,降幅21.7%,将MindDrive的平均延迟从443ms降至306.2ms,降幅30.9%;对于Orion,估计的峰值GPU内存进一步从45GB降至29GB。在与该http URL共存的实车部署中,原生Orion因车载GPU内存预算不足无法运行,而混合版本可与完整车载栈成功运行。
英文摘要
Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure. In a full autonomous driving stack, this problem is even more pronounced: legacy vehicle platforms were provisioned for modular pipelines, yet after several planning-related functions are absorbed into a unified VLA model, part of the original CPU budget becomes underutilized, while the visual encoder and the main reasoning path still concentrate most computation and memory demand on the GPU. As a result, directly deploying VLA together with the rest of the onboard system can be hard under realistic GPU memory constraints. To address this issue, we present a hybrid CPU--GPU inference framework with flexible resource scheduling for autonomous driving. Our design partitions the VLA backbone at the block-layer granularity, executes the visual encoder and LLM prefix on the GPU, and offloads the LLM suffix to the CPU through a cross-frame asynchronous pipeline, thereby exposing a schedulable boundary for redistributing compute and memory pressure across heterogeneous processors. We evaluate the proposed framework on two representative driving VLA models, Orion and MindDrive. On Bench2Drive, our method reduces average latency from 521ms to 408.0ms for Orion and from 443ms to 306.2ms for MindDrive, corresponding to 21.7% and 30.9% reduction, respectively. For Orion, the estimated peak GPU memory is further reduced from 45GB to 29GB. In real-vehicle deployment under coexistence with Autoware.Universe, native Orion cannot run because the onboard GPU memory budget is insufficient, whereas the hybrid version runs successfully together with the full vehicle stack.
发表机构
- City University of Hong Kong(香港城市大学)
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。