EcoVLA:面向实时约束的视觉-语言-动作模型的节能设备-边缘协同推理
EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints
浏览论文内容
中文总结 AI 辅助
EcoVLA是面向VLA模型的自适应设备-边缘协同推理框架,可在实时约束下最大化能效,提升能效达236%,同时保持服务水平目标满足。
中文摘要 AI 辅助
视觉-语言-动作(Vision-Language-Action,VLA)模型已成为具身智能领域极具潜力的基础模型,但其高昂的推理成本给机器人系统部署带来了重大挑战。实际应用中,设备端推理受限于有限的计算能力和能耗预算,难以同时满足实时控制和能效要求;而将推理工作负载卸载至边缘服务器则易受系统状况波动影响,引入不可预测的延迟风险。设备-边缘协同推理是一种有前景的解决方案,但针对VLA模型的系统研究仍较为匮乏,尤其是能同时解决实时约束和系统级能效的统一协同推理框架。为此,本文提出EcoVLA,这是一种面向VLA模型的自适应设备-边缘协同推理框架,可在实时约束下最大化系统能效。EcoVLA首先对不同VLA范式引入统一的阶段级抽象,建立了与架构无关的协同推理设计空间;随后构建了设备-边缘-网络联合延迟与能耗预测模型,以实现候选协同推理方案的快速运行时评估。在此基础上,EcoVLA持续选择满足实时约束的能效最优方案,仅产生毫秒级开销,可自适应网络和系统状态的运行时变化。此外,EcoVLA还采用了轻量级传输机制处理阶段间中间张量,以减少跨设备协作带来的通信开销。针对VLA模型的实验结果表明,在20Hz动作输出频率约束下,EcoVLA相比现有协同推理方法可将系统能效提升高达236%,同时在动态网络和边缘工作负载条件下始终保持服务水平目标(SLO)的满足。
英文摘要
Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inference cost poses significant challenges for deployment in robotic systems. In practice, on-device inference is constrained by limited compute capacity and energy budgets, struggling to simultaneously satisfy real-time control and energy efficiency requirements. Alternatively, offloading the inference workload to an edge server is susceptible to fluctuations in system conditions, introducing unpredictable latency risks. Device-edge co-inference offers a promising solution, but systematic research tailored to VLA models remains scarce, particularly a unified co-inference framework that jointly addresses real-time constraints and system-level energy efficiency. Thus, we propose EcoVLA, an adaptive device-edge co-inference framework for VLA models that maximizes system energy efficiency under real-time constraints. EcoVLA first introduces a unified stage-level abstraction over different VLA paradigms, establishing an architecture-agnostic co-inference design space. It then formulates a joint device-edge-network latency and energy prediction model to enable rapid runtime evaluation of candidate co-inference schemes. Building on this, EcoVLA continuously selects the energy-optimal scheme satisfying real-time constraints with millisecond-level overhead, adapting to runtime variations in network and system states. Furthermore, EcoVLA incorporates a lightweight transmission mechanism for inter-stage intermediate tensors to reduce the communication overhead incurred by cross-device collaboration. Experimental results across VLA models show that EcoVLA improves system energy efficiency by up to 236% over existing co-inference approaches under a 20 Hz action output frequency constraint, while consistently maintaining SLO satisfaction under dynamic network and edge workload conditions.
发表机构
- Beihang University(北京航空航天大学)
- Beijing University of Technology(北京工业大学)
机构由 AI 辅助整理,请以论文原文为准。