并非所有视觉 token 都可安全移除:后果敏感型视觉 token 压缩
Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression
浏览论文内容
中文总结 AI 辅助
该研究针对视觉语言模型提出后果敏感型视觉 token 压缩方法,可在相同总 token 预算下降低高风险错误,还能减少代价加权错误并降低延迟,泛化性良好。
中文摘要 AI 辅助
视觉语言模型(VLM)的视觉 token 压缩大多依赖注意力、冗余度和不确定性等标准,以在固定计算预算下最大化平均准确率,隐含假设所有错误的代价相同。然而,下游任务中错误预测的后果极少对称:误读发票金额的代价远高于误分类背景颜色。受此启发,我们提出后果敏感型视觉 token 压缩,根据请求的潜在错误代价分配视觉计算。我们的方法遵循「校准后分配」流程:离线估计特定后果的错误预算曲线,在线利用问题或任务信息提供的后果信号应用校准后的 token 预算。在受控的任务内基准测试中,高代价和低代价问题来自同一文档图像,仅靠内容无法判断哪些问题出错代价高昂。在此设置下,我们的方法在相同总 token 预算下将高风险错误从 0.300 降至 0.133,而内容驱动型分配器的表现与均匀分配无异。通过测量不同成本比率下错误率随 token 预算的变化,我们推导了分配前沿:当错误代价相同时,均匀分配最优;随着成本差距增大,向高代价问题转移 token 的益处愈发明显。该分配原理可良好泛化至三个密集视觉语言基准、两种预算实现机制(token 删除和分辨率重新分配)、两种 VLM 架构及多种 token 选择策略。在真实混合工作负载中,后果敏感型分配将代价加权错误降低 38%,同时比全分辨率推理的延迟低约 21%。
英文摘要
Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost. However, the consequence of an incorrect prediction on downstream tasks is rarely symmetric: misreading an invoice amount can be far more costly than misclassifying a background color. Motivated by this, we introduce consequence-sensitive visual token compression, which allocates visual computation across requests according to their potential error costs. Our method follows a calibrate-then-allocate procedure, estimating consequence-specific error-budget curves offline and applying the calibrated token budgets online using consequence signals available from question or task information. On a controlled within-task benchmark, high- and low-consequence questions are drawn from the same document images, so content alone cannot reveal which questions are costly to get wrong. In this setting, our method reduces high-stakes errors from 0.300 to 0.133 under the same total token budget, whereas a content-driven allocator performs no better than uniform allocation. Measuring how error rates change with token budget across different cost ratios, we derive an allocation frontier: uniform allocation is optimal when errors are equally costly, and token transfer toward high-consequence questions becomes increasingly beneficial as the cost gap grows. This allocation principle generalizes well across three dense vision-language benchmarks, two budget realization mechanisms (token deletion and resolution reallocation), two VLM architectures, and multiple token selection strategies. On a realistic mixed workload, consequence-sensitive allocation reduces cost-weighted error by 38% while achieving approximately 21% lower latency than full-resolution inference.
发表机构
- The University of Sydney(悉尼大学)
- Tongji University(同济大学)
- University of Surrey(萨里大学)
- Nankai University(南开大学)
- Cornell University(康奈尔大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。