发表机构
Jeonbuk National University; Korea Institute of Industrial Technology (KITECH); Chung-Ang University(全北国立大学; 韩国生产技术研究院; 中央大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReFIT框架通过指令引导的窗口重塑与令牌细化,在降低计算成本的同时提升视觉问答准确性。
AI 中文摘要
大型视觉语言模型在视觉问答任务上取得了强劲性能,但处理高分辨率、信息丰富的图像需要大量计算,这促使了视觉令牌缩减。然而,现有方法常常修剪单个令牌或依赖固定尺寸裁剪,限制了它们保留空间结构信息(如水平或垂直延伸的文本)的能力。为解决这一局限,我们提出了ReFIT,一种用于高效大型视觉语言模型推理的指令引导视觉令牌缩减框架。ReFIT由相关性引导窗口重塑(RWR)和指令引导令牌细化(ITR)组成,其中RWR通过适应指令相关区域的空间特征来捕获这些区域,而ITR进一步移除不必要的视觉令牌。在四个视觉问答基准上的实验表明,ReFIT在降低计算成本的同时提高了答案准确性,定性结果也证明了其在定位相关区域和移除不必要视觉信息方面的有效性。
英文摘要
Large Vision-Language Models achieve strong VQA performance, but processing high-resolution, information-rich images requires substantial computation, motivating visual token reduction. However, existing methods often prune individual tokens or rely on fixed-size cropping, limiting their ability to preserve spatially structured information such as horizontally or vertically elongated text. To address this limitation, we propose ReFIT, an instruction-guided visual token reduction framework for efficient LVLM inference. ReFIT consists of Relevance-Guided Window Reshaping (RWR) and Instruction-Guided Token Refinement (ITR), where RWR captures instruction-relevant regions by adapting to their spatial characteristics, while ITR further removes unnecessary visual tokens. Experiments on four VQA benchmarks demonstrate that ReFIT improves answer accuracy while reducing computational cost, and qualitative results demonstrate its effectiveness in localizing relevant regions and removing unnecessary visual information.
Comments5 pages, 2 figures. Under review