AI 中文总结
FOVEA是一种缓存友好型多模态推测解码方法,通过动态检索视觉记忆子集提升草稿接受率,使多模态生成速度最高提升2.13倍。
AI 中文摘要
多模态推测解码通过让轻量级草稿模型提出候选token,供更大规模的目标模型并行验证,以此加速视觉-语言模型。现有方法通常将草稿模型的条件设定在固定视觉接口上,例如预定义的视觉token预算或静态压缩表示。然而,我们开展的可控视觉预算分析显示,不同任务和解码阶段的视觉需求存在显著差异,这意味着更多视觉输入并不总是有益的。实际上,证据不足可能会削弱视觉定位能力,而过多上下文则会增加开销并可能干扰草稿生成。我们提出FOVEA(Focused On-demand Visual Evidence Adaptation),一种缓存友好型方法,该方法构建可复用的视觉记忆,并为草稿状态动态检索有界子集。累积质量规则决定了选择的条目数量和具体条目。选中的条目被聚合成视觉读数,并通过轻量级门控残差校正与当前草稿隐藏状态融合。该校正不会将视觉token插入自回归上下文,仅修改传递给语言模型头部的表示。在多个视觉-语言骨干网络和多模态基准上开展的实验表明,FOVEA提高了草稿接受率和端到端解码速度,实现了比自回归解码最高达2.13倍的加速。这些结果证明,基于状态的证据检索是多模态生成中复用固定视觉表示的有效替代方案。
英文摘要
Multimodal speculative decoding accelerates vision-language models by allowing a lightweight draft model to propose candidate tokens for parallel verification by a larger target model. Existing methods typically condition the drafter on a fixed visual interface, such as a predefined visual-token budget or a static compressed representation. However, our controlled visual-budget analysis shows that visual demand varies substantially across tasks and decoding stages, which means more visual input is not always beneficial. Actually, insufficient evidence may weaken visual grounding, while excessive context adds overhead and may disrupt drafting. We propose FOVEA (Focused On-demand Visual Evidence Adaptation), a cache-friendly approach that builds a reusable visual memory and dynamically retrieves a bounded subset for a draft state. A cumulative-mass rule determines both how many and which entries are selected. The selected entries are aggregated into a visual readout and fused with the current draft hidden state through a lightweight gated residual correction. Rather than inserting visual tokens into the autoregressive context, the correction modifies only the representation passed to the language-model head. Experiments across multiple vision-language backbones and multimodal benchmarks show that FOVEA improves draft acceptance and end-to-end decoding speed, achieving up to $2.13\times$ speedup over autoregressive decoding. These results demonstrate that state-conditioned evidence retrieval is an effective alternative to reusing a fixed visual representation throughout multimodal generation.