TIGER:用于多模态推测解码的文本条件视觉门控路由与接受对齐
TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding
浏览论文内容
中文总结 AI 辅助
针对多模态推测解码中草稿模型视觉关键内容易偏差及现有方法不足的问题,提出TIGER框架,基于文本状态动态选视觉令牌,用接受对齐分组策略训练优化草稿模型,实验证明其在多方面取得良好效果。
中文摘要 AI 辅助
推测解码通过让轻量级草稿模型提出多个由更大的目标模型验证的令牌来加速自回归生成。虽然对仅文本的语言模型有效,但在视觉语言模型中收益有限,因为草稿模型在视觉关键内容上常出现偏差,且现有多模态加速方法未直接解决无关视觉证据或优化验证器接受的前缀长度。我们提出了TIGER,一种用于多模态推测解码的文本条件视觉门控路由框架。TIGER基于草稿模型的当前文本状态动态选择一组与上下文相关的稀疏视觉令牌。为使训练与推理效率更好对齐,我们使用基于验证器推导奖励的接受对齐分组策略训练来优化草稿模型,该奖励基于接受的前缀长度,并基于KL锚定的蒸馏热启动构建。实验表明,TIGER在精确验证器侧推测解码下,在接受前缀长度和推测加速方面取得了一致的收益,同时在视觉路由分析中以可比的下游精度实现了良好的质量-延迟权衡。
英文摘要
Speculative decoding accelerates autoregressive generation by letting a lightweight drafter propose multiple tokens that are verified by a larger target model. Although effective for text-only LLMs, speculative decoding yields limited gains in VLMs because drafters often diverge on vision-critical content, while existing multimodal acceleration methods do not directly address irrelevant visual evidence or optimize the verifier-accepted prefix length that governs speedup. We propose TIGER, a Text-conditioned vIsual GatEd Routing framework for multimodal speculative decoding. TIGER dynamically selects a sparse set of context-relevant visual tokens based on the drafter's current textual state, rather than expose the full visual token set or a fixed compressed interface. To better align training with inference-time efficiency, we optimize the drafter with acceptance-aligned group-based policy training using verifier-derived rewards based on accepted prefix length, built on top of distillation warm start with KL anchoring. This encourages the drafter not only to imitate the target model, but also to produce speculative continuations that survive verification for longer prefixes. Experiments show that TIGER yields consistent gains in accepted prefix length and speculative speedup under exact verifier-side speculative decoding, while achieving favorable quality-latency trade-offs with comparable downstream accuracy in visual-routing analyses.
发表机构
- National University of Singapore(新加坡国立大学)
- Nanyang Technological University(南洋理工大学)
- Center of AI Research, VinUniversity(Vin大学人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。