arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

眼见无需耗能,言语却要代价:揭示边缘视觉语言模型推理中的真正能量瓶颈

Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference

Junfei Zhan, Haoxun Shen, Mingang Guo, Zixuan Huang, Tengjiao He

arXiv 2607.09520首次发表:更新:

发表机构

University of Pennsylvania; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; Jinan University(宾夕法尼亚大学; 中国科学院深圳先进技术研究院; 暨南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究边缘VLM推理的能量瓶颈,通过系统能量分析发现平均推理功率恒定,输出令牌耗时多是关键,图像复杂度因输出长度影响能量,揭示视觉令牌修剪局限,控制输出长度能大幅节能。

AI 中文摘要

视觉语言模型(VLM)是具身人工智能的感知支柱,但其在边缘硬件上的能源足迹仍未得到充分理解。现有提高效率的努力主要集中在减少视觉令牌上,隐含地将视觉处理视为主要的能源成本。我们通过对设备上VLM推理进行首次系统的能量分析,推翻了这一隐含假设,该分析涵盖了三个架构家族的五个模型、四种输入分辨率和两个硬件平台(NVIDIA RTX 3070和Jetson Orin NX)。我们的分析得出了三个发现。首先,平均推理功率是一个模型内在常数,与输入分辨率、图像复杂度和提示类型无关,在所有条件下变化小于5%。这意味着所有输入之间的能量变化必须源于推理时间的变化,而不是功耗的变化。其次,由于预填充和解码之间的计算和内存不对称,每个输出令牌的挂钟时间比每个输入令牌多11到39倍,使得输出令牌数量成为延迟和能量的主要驱动因素。第三,以图像中的对象数量衡量的图像复杂度,在相同分辨率下会导致高达4.1倍的能量差异。这种变化不是源于视觉处理成本的增加,而是源于输出长度的差异。这些发现揭示了视觉令牌修剪的一个基本局限性:即使移除所有视觉令牌,对于固定令牌模型最多也只能节省10%的总能量。在跨越10亿到80亿参数的模型中,控制输出长度最多可节省97%的总能量,并且在更大的模型规模下,解码的能量主导地位会变得更强。简而言之,边缘VLM推理中的真正能量瓶颈不是模型看到了什么,而是它说了多少。

英文摘要

Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly understood. Existing efficiency efforts focus predominantly on reducing visual tokens, implicitly treating visual processing as the dominant energy cost. We overturn this implicit assumption through the first systematic energy profiling of on-device VLM inference, spanning five models across three architecture families, four input resolutions, and two hardware platforms (NVIDIA RTX 3070 and Jetson Orin NX). Our analysis yields three findings. First, average inference power is a model-intrinsic constant, invariant to input resolution, image complexity, and prompt type, with less than 5% variation across all conditions. This means that all energy variation across inputs must arise from variation in inference time, not from variation in power draw. Second, each output token costs 11 to 39x more wall-clock time than each input token due to the compute-bound and memory-bound asymmetry between prefill and decode, making output token count the dominant driver of both latency and energy. Third, image complexity, measured by the number of objects in an image, induces up to 4.1x energy differences at identical resolution. This variation arises not from increased visual processing cost, but from differences in output length. These findings expose a fundamental limitation of visual token pruning: even removing all visual tokens saves at most 10% of total energy for fixed-token models. Across models spanning 1 billion to 8 billion parameters, controlling output length saves up to 97% of total energy, with the energy dominance of decoding growing stronger at larger model scale. In short, the true energy bottleneck in edge VLM inference is not what the model sees, but how much it says. Code is available at https://github.com/Junfei-Z/seeing-is-free.

CommentsAccepted to ACM MM 2026. This version includes the appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑