利用卫星图像中答案不变冗余实现边缘端高效VLM推理
Exploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge
浏览论文内容
中文总结 AI 辅助
针对卫星图像边缘端VLM推理,提出利用答案不变冗余的Rift系统,通过查询条件分块剪枝与弹性预填充,在Jetson AGX Orin上降低能耗78%、延迟69%,准确率提升至73%。
中文摘要 AI 辅助
星载视觉语言模型可使卫星直接回答查询,但对高分辨率图像进行穷举分块推理既缓慢又耗能。我们识别出答案不变令牌冗余(AITR):即在不改变最终答案的前提下可移除的图像块和视觉令牌。我们提出Rift,一种两阶段系统,先执行查询条件分块剪枝,再进行弹性预填充以减少令牌预算。我们在运行于Jetson AGX Orin上的LLaVA-1.5 7B上评估该系统。与穷举分块推理相比,Rift将能耗降低78%,延迟降低69%,同时将准确率从45%提升至73%。
英文摘要
Onboard vision-language models could enable satellites to answer queries directly, but exhaustive tiled inference over high-resolution imagery is slow and energy-intensive. We identify answer-invariant token redundancy (AITR): image tiles and vision tokens that can be removed without changing the final answer. We present Rift, a two-stage system that performs query-conditioned tile pruning followed by elastic prefill to reduce token budget. We evaluate it on LLaVA-1.5 7B running on Jetson AGX Orin. Compared with exhaustive tiled inference, Rift reduces energy by 78% and latency by 69%, while increasing accuracy from 45% to 73%.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。