arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ET-Prune:面向文本丰富型多模态大语言模型的视觉 token 剪枝的证据感知动态预算分配

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Zizhong Ding, Junxian Li, Kai Liu, Shaoqiu Zhang, Xiao Xiao, Linghe Kong, Yulun Zhang

arXiv 2608.01979首次发表:更新:

AI 中文总结

ET-Prune 是一种无需训练的视觉 token 剪枝框架,通过证据分配实现动态预算,在文本丰富的多模态任务中,于约一半 token 保留量下,在 OCRBench-v2、MMBench v1.1 等基准上取得优于基线的性能,实现了质量与成本的良好平衡。

AI 中文摘要

视觉 token 剪枝可降低多模态大语言模型的推理成本,但固定 token 比例与文本丰富的输入匹配度较差。在以 OCR 为核心的任务中,决定性证据可能是问题指定的少量字符、标签或字段,无差别剪枝可能会抹去这些证据,同时保留视觉上显著但不相关的区域。本文提出 ET-Prune,这是一种无需训练的框架,将剪枝转化为证据分配问题。它从解码器侧的部分查询-键块中推导问题条件证据,保护类文本的空间区域,并将证据不确定性和密度转换为样本特定的 token 下限。随后,三个渐进式中间层事件使序列向该预算调整,为分散或文本密集的证据保留更多 token,对集中的证据则更激进地剪枝。在每个配置进行一次确定性推理得到的观测点估计下,ET-Prune 在约一半 token 数量下,于所有六个骨干-基准比较中领先或与其他剪枝方法持平。在 OCRBench-v2 上,它在 Qwen3-VL-8B 和 InternVL3.5-8B 上分别比最强的剪枝基线领先 1.80 和 0.68 个百分点,同时保留约一半的视觉 token;在 MMBench v1.1 上,当平均视觉 token 保留率为 54.45%时,其循环精确匹配准确率达到 0.8467,而 Vanilla 为 0.8437。这些结果表明,在文本丰富的多模态推理中,证据感知动态预算分配具有良好的观测质量-成本权衡。

英文摘要

Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In OCR-centric tasks, decisive evidence can be a small number, label, or field whose relevance is specified by the question; indiscriminate pruning can erase that evidence while retaining visually salient but irrelevant regions. We present ET-Prune, a training-free framework that casts pruning as evidence allocation. It derives question-conditioned evidence from a decoder-side partial query-key block, safeguards text-like spatial regions, and converts evidence uncertainty and density into a sample-specific token floor. Three progressive middle-layer events then move the sequence toward this budget, retaining more tokens for diffuse or text-dense evidence and pruning concentrated evidence more aggressively. At the observed point estimates from one deterministic pass per configuration, ET-Prune leads or ties among pruned methods in all six backbone-benchmark comparisons at roughly half tokens. On OCRBench-v2, it leads the strongest pruned baselines by 1.80 and 0.68 percentage points on Qwen3-VL-8B and InternVL3.5-8B, respectively, while retaining about half of the visual tokens; on MMBench v1.1, it reaches 0.8467 circular exact-matching accuracy versus 0.8437 for Vanilla at 54.45% average visual-token retention. These results show a favorable observed quality-cost trade-off for evidence-aware dynamic budgeting in text-rich multimodal inference.

CommentsCode and supplementary material is at https://github.com/Labyrinth0419/ET-Prune

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑