arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29029cs.CVcs.LG

利用卫星图像中答案不变冗余实现边缘端高效VLM推理

Exploiting answer-invariant redundancies in satellite imagery for efficient VLM inference on edge

Ishani Janveja, Davis Zhang, Seoyul Oh, Deepak Vasisht

首次发表
浏览论文内容

中文总结 AI 辅助

针对卫星图像边缘端VLM推理,提出利用答案不变冗余的Rift系统,通过查询条件分块剪枝与弹性预填充,在Jetson AGX Orin上降低能耗78%、延迟69%,准确率提升至73%。

中文摘要 AI 辅助

星载视觉语言模型可使卫星直接回答查询,但对高分辨率图像进行穷举分块推理既缓慢又耗能。我们识别出答案不变令牌冗余(AITR):即在不改变最终答案的前提下可移除的图像块和视觉令牌。我们提出Rift,一种两阶段系统,先执行查询条件分块剪枝,再进行弹性预填充以减少令牌预算。我们在运行于Jetson AGX Orin上的LLaVA-1.5 7B上评估该系统。与穷举分块推理相比,Rift将能耗降低78%,延迟降低69%,同时将准确率从45%提升至73%。

英文摘要

Onboard vision-language models could enable satellites to answer queries directly, but exhaustive tiled inference over high-resolution imagery is slow and energy-intensive. We identify answer-invariant token redundancy (AITR): image tiles and vision tokens that can be removed without changing the final answer. We present Rift, a two-stage system that performs query-conditioned tile pruning followed by elastic prefill to reduce token budget. We evaluate it on LLaVA-1.5 7B running on Jetson AGX Orin. Compared with exhaustive tiled inference, Rift reduces energy by 78% and latency by 69%, while increasing accuracy from 45% to 73%.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑