arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09176cs.CVcs.AI

并非所有视觉 token 都可安全移除:后果敏感型视觉 token 压缩

Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression

Jingbo Wen, Liang He, Mingyu Cao, Haoyu Wang, Minxuan Hu, Kangning Cui, Xilu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对视觉语言模型提出后果敏感型视觉 token 压缩方法,可在相同总 token 预算下降低高风险错误,还能减少代价加权错误并降低延迟,泛化性良好。

中文摘要 AI 辅助

视觉语言模型(VLM)的视觉 token 压缩大多依赖注意力、冗余度和不确定性等标准,以在固定计算预算下最大化平均准确率,隐含假设所有错误的代价相同。然而,下游任务中错误预测的后果极少对称:误读发票金额的代价远高于误分类背景颜色。受此启发,我们提出后果敏感型视觉 token 压缩,根据请求的潜在错误代价分配视觉计算。我们的方法遵循「校准后分配」流程:离线估计特定后果的错误预算曲线,在线利用问题或任务信息提供的后果信号应用校准后的 token 预算。在受控的任务内基准测试中,高代价和低代价问题来自同一文档图像,仅靠内容无法判断哪些问题出错代价高昂。在此设置下,我们的方法在相同总 token 预算下将高风险错误从 0.300 降至 0.133,而内容驱动型分配器的表现与均匀分配无异。通过测量不同成本比率下错误率随 token 预算的变化,我们推导了分配前沿:当错误代价相同时,均匀分配最优;随着成本差距增大,向高代价问题转移 token 的益处愈发明显。该分配原理可良好泛化至三个密集视觉语言基准、两种预算实现机制(token 删除和分辨率重新分配)、两种 VLM 架构及多种 token 选择策略。在真实混合工作负载中,后果敏感型分配将代价加权错误降低 38%,同时比全分辨率推理的延迟低约 21%。

英文摘要

Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost. However, the consequence of an incorrect prediction on downstream tasks is rarely symmetric: misreading an invoice amount can be far more costly than misclassifying a background color. Motivated by this, we introduce consequence-sensitive visual token compression, which allocates visual computation across requests according to their potential error costs. Our method follows a calibrate-then-allocate procedure, estimating consequence-specific error-budget curves offline and applying the calibrated token budgets online using consequence signals available from question or task information. On a controlled within-task benchmark, high- and low-consequence questions are drawn from the same document images, so content alone cannot reveal which questions are costly to get wrong. In this setting, our method reduces high-stakes errors from 0.300 to 0.133 under the same total token budget, whereas a content-driven allocator performs no better than uniform allocation. Measuring how error rates change with token budget across different cost ratios, we derive an allocation frontier: uniform allocation is optimal when errors are equally costly, and token transfer toward high-consequence questions becomes increasingly beneficial as the cost gap grows. This allocation principle generalizes well across three dense vision-language benchmarks, two budget realization mechanisms (token deletion and resolution reallocation), two VLM architectures, and multiple token selection strategies. On a realistic mixed workload, consequence-sensitive allocation reduces cost-weighted error by 38% while achieving approximately 21% lower latency than full-resolution inference.

发表机构

  • The University of Sydney(悉尼大学)
  • Tongji University(同济大学)
  • University of Surrey(萨里大学)
  • Nankai University(南开大学)
  • Cornell University(康奈尔大学)
  • City University of Hong Kong(香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

↑