发表机构
University of Illinois at Urbana-Champaign; Microsoft Research; Google DeepMind(伊利诺伊大学厄巴纳-香槟分校; 微软研究院; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出轻量级ReToken,通过选择视觉键值缓存中与查询相关的稀疏标记解决长视觉上下文的视觉检索难题,在多基准测试中提升Qwen3VL-8B等模型性能,且可在单H100运行。
AI 中文摘要
长视觉上下文对视觉-语言模型构成挑战:随着干扰项数量增加,模型性能会下降,且在GPU内存限制下一次性处理所有标记在计算上不可行。我们提出ReToken,这是一个作为显式检索目标训练的单一可学习嵌入,用于从预填充的视觉键值缓存中选择与查询相关的稀疏视觉标记。仅在小型图像-问答数据集上训练的ReToken,在图像和视频基准测试中均取得稳定提升:在Visual Haystacks上,它将Qwen3VL-8B提升13.4个点,将InternVL3.5提升12.4个点(相对提升超20%);在LVBench上,它零样本迁移到长视频场景,为Qwen3VL-8B带来8.0个点的提升。得益于其轻量级设计,训练和长视频推理均可在单个H100上完成。代码可在该https URL获取。
英文摘要
Long visual contexts challenge vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once can exceed GPU memory limits. We present RETOKEN, a single learnable embedding that extracts retrieval signals from the VLM's internal representations to select query-relevant visual tokens from the pre-filled KV cache. This enables retrieval within the answering VLM, without a separate retriever or re-encoding. Despite being trained on only a small image-QA dataset, RETOKEN generalizes across image and video benchmarks. On Visual Haystacks, it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (>20% relative). On LVBench, it transfers zero-shot to long video and improves Qwen3VL-8B by 8.0 points. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at: https://github.com/avaxiao/ReToken
CommentsAccepted to NeurIPS 2026. Code: https://github.com/avaxiao/ReToken