AI 中文总结
该研究通过可复现协议分析视觉令牌数量与多模态推理延迟的盈亏平衡点,发现视觉前路由在A100上的延迟优势超过视觉后策略的下游令牌减少量,且视觉后预测器的效果经校正后仍显著。
AI 中文摘要
更少的视觉令牌并不一定能保证更低的端到端延迟。我们采用可复现的协议评估盈亏平衡点,该协议考虑了决策开销、共享工作以及各策略可避免的算子。阶段级分解将这些组件与实测的端到端延迟相协调。在包含30个样本的试点研究中,尽管存在状态复用,两种被测的自回归探测器仍比全量模型更慢。轻量型视觉后预测器在RTX 3090和A100上产生的配对置信区间低于零,且经保守的全对Holm校正后仍具显著性。视觉前图像尺寸规则在两款GPU上同样产生低于零的区间,但经相同校正后无显著性。视觉前路由具备视觉后剪枝所不具备的结构优势:它可避免预处理和视觉编码。在A100上,该优势超过了视觉后策略近8倍的下游令牌减少量。报告的质量以全量模型正确回答的样本为条件,并非基准准确率。
英文摘要
Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operators each policy can avoid. A stage-level decomposition reconciles these components with measured end-to-end latency. In a 30-example pilot, the two tested autoregressive probes remain slower than Full despite state reuse. A lightweight post-vision predictor yields paired confidence intervals below zero on RTX 3090 and A100 and remains significant after a conservative all-pairs Holm correction. A pre-vision image-size rule also yields intervals below zero on both GPUs, although neither comparison remains significant after the same correction. Pre-vision routing has a structural opportunity unavailable to post-vision pruning: it can avoid preprocessing and vision encoding. On A100, this opportunity outweighs a nearly eightfold larger downstream token reduction by the post-vision policy. Reported quality is conditional on examples answered correctly by Full and is not benchmark accuracy.
Comments16 pages, 3 figures, 13 tables. Experiments use Qwen2.5-VL-3B-Instruct on RTX 3090 and A100 PCIe GPUs