arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04132cs.CV

RUTA:基于率-效用优化的原则性视觉令牌分配方法

RUTA: Principled Visual Token Allocation via Rate-Utility Optimization

  • City University of Hong Kong(香港城市大学)
  • Google Inc.(谷歌公司)

机构由 AI 辅助整理,请以论文原文为准。

Jian Zou, Xiaoyu Xu, Zhihua Wang, Yilin Wang, Balu Adsumilli, Kede Ma

AI总结:

针对视觉语言模型因长视觉令牌序列导致的LLM侧计算和内存成本过高问题,提出RUTA方法,仅用2.0%、4.2%的视觉令牌,分别在LLaVA-NeXT-7B、Qwen3-VL-8B上保留88.2%、94.4%的任务性能。

AI中文摘要:

高分辨率图像和长视频为视觉语言模型提供了用于多模态推理和细粒度感知的丰富上下文,但由此产生的长视觉令牌序列会导致大语言模型侧的计算和内存成本过高。现有视觉令牌缩减器通常以规定的速率运行,而近期的方法会使用特定于方法的学习阈值或重要性预测器来跨输入调整令牌数量。我们提出了RUTA,这是一种原则性的率-效用令牌分配方法,在进入大语言模型(LLM)之前执行缩减,同时学习要保留的令牌以及为每个图像-查询对分配多少令牌。RUTA构建查询条件的候选令牌,并为每个候选令牌预测保留概率。在训练期间,这些概率参数化独立的伯努利门,而它们的总和提供了每对令牌数量的可微分训练时估计值。保留的令牌充当锚点,根据语义亲和性和空间邻近性聚合来自未保留令牌的信息。RUTA通过惩罚性的率-效用目标进行优化,该目标平衡下游任务损失与预期令牌使用量。在五个基准上取平均值,并相对于每个主干的全令牌基线进行测量,RUTA仅使用2.0%和4.2%的视觉令牌,同时在LLaVA-NeXT-7B和Qwen3-VL-8B上分别保留了88.2%和94.4%的任务性能。

英文摘要:

High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the resulting long visual token sequences make large language model-side computation and memory costly. Existing visual token reducers often operate at prescribed rates, while recent methods adapt token counts across inputs using method-specific learned thresholds or importance predictors. We introduce RUTA, a principled Rate-Utility Token Allocation method that performs pre-LLM reduction by jointly learning which tokens to retain and how many to allocate to each image-query pair. RUTA constructs query-conditioned candidate tokens and predicts a retention probability for each candidate. During training, these probabilities parameterize independent Bernoulli gates, while their sum provides a differentiable training-time estimate of the token count for each pair. Retained tokens serve as anchors that aggregate information from non-retained tokens according to semantic affinity and spatial proximity. RUTA is optimized with a penalized rate-utility objective that balances downstream task loss against expected token usage. Averaged across five benchmarks and measured relative to each backbone's full-token baseline, RUTA uses only $2.0\%$ and $4.2\%$ of visual tokens while preserving $88.2\%$ and $94.4\%$ of task performance on LLaVA-NeXT-7B and Qwen3-VL-8B, respectively.

↑