结合推测推理与验证的高效VLA推理的算法-架构协同设计
Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification
- School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
- Computer Science, King Abdullah University of Science and Technology(阿卜杜拉国王科技大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对VLA模型推理延迟高、动作长度有限的问题,提出SpecVLA算法-架构协同框架,通过状态感知推理、sVLA验证模型与异构架构,在保持任务成功率的同时降低延迟,实现高效可靠的实时机器人操控。
AI中文摘要:
视觉-语言-动作(Vision-Language-Action,VLA)模型在具身人工智能领域展现出卓越能力,但其高计算成本与有限的预测动作长度阻碍了实时部署。尽管专为高效具身人工智能设计的专用加速器Dadu-Corkki已被提出,但该加速器未利用机器人与其环境之间的固有交互模式,导致预测动作长度相对较短。我们观察到机器人环境自然在活跃状态(此时精确动作至关重要)与非活跃状态(此时动作对任务成功的影响有限)之间交替。这一发现提供了新的调度机会:在非活跃状态下进行长动作长度的推测预测,同时在活跃状态下进行选择性验证。我们提出SpecVLA,这是一个算法-系统协同设计框架,可自适应平衡动作长度、推理延迟与任务可靠性。在算法层面,SpecVLA引入了状态感知的VLA推理执行范式,以及利用差分残差和分块混合精度量化构建的硬件友好型小型验证模型(sVLA)。在系统层面,我们开发了由GPU与机器人专用硬件模块组成的异构架构,以及通过并行执行解耦VLA与sVLA的推测数据流。在OpenVLA和RDT上针对LIBERO与ManiSkill基准的综合评估显示,SpecVLA在保持任务成功率的同时显著降低了端到端延迟。通过实现带及时验证的长动作长度推测预测,SpecVLA达成了兼具高效率与高可靠性的实时机器人操控。
英文摘要:
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cost and limited predicted action length hinder real-time deployment. Although Dadu-Corki, a dedicated accelerator for efficient embodied AI, has been introduced, it does not exploit the inherent interaction patterns between the robot and its environment, which results in a relatively short predicted action length. We observe that robotic environments naturally alternate between active states-where precise actions are crucial-and inactive states-where actions have limited impact on task success. This insight enables a new scheduling opportunity: long-action-length speculative prediction in inactive states, paired with selective verification in active states. We propose SpecVLA, an algorithm-system co-design framework that adaptively balances action length, inference latency, and task reliability. On the algorithm side, SpecVLA introduces a state-aware VLA inference execution paradigm and a hardware-friendly construction of a smaller verification model (sVLA) using differential residuals and block-wise mixed-precision quantization. On the system side, we develop a heterogeneous architecture consisting of a GPU and a robotic-specific hardware module, along with a speculative dataflow that decouples VLA and sVLA through parallel execution. Comprehensive evaluations on OpenVLA and RDT across LIBERO and ManiSkill benchmarks show that SpecVLA reduces end-to-end latency significantly while preserving task success rate. By enabling long-action-length speculative prediction with timely verification, SpecVLA achieves real-time robotic manipulation with both high efficiency and reliability.