arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiffPrune:用于视觉-语言模型中 token 剪枝的可微信息节流方法

DiffPrune: differentiable information throttling for token pruning in vision-language models

Landi He, Mingde Yao, Shawn Young, Lijian Xu

arXiv 2608.01985首次发表:更新:

发表机构

Shenzhen University of Advanced Technology; CUHK MMLab, CPII under InnoHK(深圳先进技术大学; 香港中文大学MMLab,InnoHK下属CPII)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DiffPrune 是一种用于视觉-语言模型的可微 token 剪枝方法,通过信息节流器避免松弛选择的不稳定性,在保留 96.5% 准确率的同时实现 2.85 倍 LLM 预填充加速,推理开销仅 0.69 毫秒。

AI 中文摘要

视觉 token 剪枝通过移除冗余视觉 token 降低视觉-语言模型(VLMs)的计算成本,核心是学习衡量 token 有用性的分数。现有方法通常依赖 Gumbel-Softmax 在训练期间近似离散选择,这类选择器使分数依赖于松弛剪枝算子的行为,而非信息损失的结果。本文提出 DiffPrune,赋予 token 分数直接意义:训练时保留所有 token,根据分数削弱每个 token 的信息,若削弱某 token 会损害任务,则推动评分器保护该 token;若不损害,则该 token 可获得更低分数。由于损失通过实际信息节流路径求导,评分器避免了松弛 token 选择的不稳定代理路径。DiffPrune 通过信息节流器实现该思路,向视觉 token 注入方差保持噪声,高分数 token 保留接近原始表示,低分数 token 携带更少原始信息;推理时移除节流器,用学习到的分数执行硬 top-K 剪枝。在 10 个 VLM 基准测试中,DiffPrune 保留了完整模型 96.5% 的准确率,同时将 LLM 预填充加速 2.85 倍,仅产生 0.69 毫秒的推理开销,代码将公开可用。

英文摘要

Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. The key is to learn a score that measures whether a token is useful. Existing methods typically rely on Gumbel-Softmax to approximate discrete selection during training. Such selectors make the score depend on the behavior of a relaxed pruning operator, not directly on the consequence of information loss. In this paper, we propose DiffPrune, which gives token scores a direct meaning. During training, DiffPrune keeps all tokens and weakens each token's information according to its score. If weakening a token hurts the task, the scorer is pushed to protect it; if not, the token can receive a lower score. Because the loss is differentiated through this actual information-throttling path, the scorer avoids the unstable surrogate path of relaxed token selection. DiffPrune implements this idea with an Information Throttler, which injects variance-preserving noise into visual tokens, where high-score tokens remain close to their original representations, while low-score tokens carry less original information. At inference, the throttler is removed, and hard top-K pruning is applied using the learned scores. Across ten VLM benchmarks, DiffPrune retains 96.5% of full-model accuracy while accelerating LLM prefill by 2.85x, with only 0.69 ms inference overhead. Code will be publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑