StackTok:通过预算自适应视觉标记选择加速视觉语言模型推理
StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection
浏览论文内容
中文总结 AI 辅助
StackTok是一种无需训练的视觉标记筛选器,通过查询相关性与预算校准的视觉覆盖动态平衡,在多个VLM和基准上以最少标记实现高推理性能。
中文摘要 AI 辅助
在视觉语言模型(VLM)中,提高图像分辨率会产生越来越长的视觉标记序列,从而大幅增加推理成本。为在不重新训练的情况下降低这一开销,现有方法会选择紧凑的标记子集,这些子集优先考虑查询相关性、视觉覆盖度或两者之间的固定权衡。然而,合适的平衡因查询和标记预算而异:局部性问题更倾向于相关性,而整体性问题则需要更广泛的视觉覆盖。我们提出StackTok,一种无需训练的筛选器,它将查询相关性作为目标,将视觉覆盖度作为预算校准的支持。StackTok从仅覆盖度的贪心序列构建一个按大小索引的覆盖参考,并使用查询-视觉亲和熵调整其支持目标。随后,一种参考门控的交替选择策略根据当前子集的支持度赤字,在相关性导向和覆盖度导向的添加之间切换。对于高分辨率输入,StackTok根据局部提名标记的合并边际增益,在多个裁剪块之间分配一个共享的标记预算。在五个VLM和十个不同的图像理解基准上的评估中,StackTok在每个测试的模型-预算设置中均位列无训练筛选器之首。在高分辨率LLaVA-NeXT-7B上,它仅使用2,880个视觉标记中的160个(5.6%)便保留了全标记性能的95.26%。
英文摘要
Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost. To reduce this overhead without retraining, existing methods select compact token subsets that prioritize query relevance, visual coverage, or a fixed trade-off between them. The appropriate balance, however, varies across queries and token budgets: localized questions favor relevance, whereas holistic questions demand broader visual coverage. We introduce StackTok, a training-free selector that treats query relevance as the objective and visual coverage as budget-calibrated support. StackTok builds a size-indexed coverage reference from a coverage-only greedy sequence and adjusts its support target using query--vision affinity entropy. A reference-gated interleaved selection policy then switches between relevance- and coverage-oriented additions according to the current subset's support deficit. For high-resolution inputs, StackTok allocates one shared token budget across crops according to the combined marginal gain of locally nominated tokens. Evaluated with five VLMs over ten distinct image-understanding benchmarks, StackTok ranks first among training-free selectors in every tested model--budget setting. On high-resolution LLaVA-NeXT-7B, it retains 95.26% of full-token performance with only 160 of 2{,}880 (5.6%) visual tokens.
发表机构
- Sichuan University(四川大学)
机构由 AI 辅助整理,请以论文原文为准。