arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉并非额外开销:视觉语言模型中用于无损投机解码的单遍块草稿方法

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models

Jungseob Lee, Seongtae Hong, Dongyub Jude Lee, Chanjun Park, Jaehyung Seo, Sugyeong Eo, Heuiseok Lim

arXiv 2609.00355首次发表:更新:

发表机构

Korea University; Zoom Communications; Soongsil University; Konkuk University; Yonsei University(高丽大学; Zoom通讯公司; 崇实大学; 建国大学; 延世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉语言模型投机解码的缺陷,提出GLANCE单遍块草稿方法,实现无损解码,解码速度较自回归最高提升2.93倍,接受块长度提升2.7倍。

AI 中文摘要

投机解码可在不改变输出的前提下加速生成,但在视觉语言模型(VLM)中,它陷入了一个适得其反的循环:草稿器保持自回归特性,因此必须保持小型;小型草稿器无法在每一步都处理图像,故视觉信息被压缩、剪枝或隐藏;脱离图像的草稿器在图像使文本可预测的区域可靠性最低。我们提出GLANCE,这是首个针对未修改的VLM目标的无损单遍块草稿方法,它从两端打破该循环:块扩散头读取目标已融合的视觉语言状态,因此视觉信息对草稿器无额外开销,且能在一次前向传播中填充整个块,故深度无需顺序步骤;宽候选树在一次目标传播中被验证,所有经审核的提示均能精确复现贪心解码。接地工作负载最受益于此,进入逐字复制机制,自回归草稿器每生成一个token需一次传播,而块草稿器仅需一次传播;在同一引擎和一轮预算下,GLANCE的解码速度比自回归快达2.93倍,其每轮仅需一次草稿传播,而生产级EAGLE3-VL头需八次,且接受的块长度是在同一语料库上训练的EAGLE-3头的2.7倍。一条规律组织这些结果:接受长度由目标的下一个token熵设定,拟合斜率随五个任务的接地程度增加而变陡;该规律可跨目标和模态迁移,并命名了自身边界,因为自由运行的文本仍倾向于链式结构。我们的代码可在该https URL获取。

英文摘要

Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target's already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree in one pass and commits exactly its greedy output. In one production engine at a fixed round budget, GLANCE decodes up to 3.05 times faster than autoregressive decoding and outpaces the production EAGLE3-VL head on average and by about 11% on grounded tasks. An entropy law explains when drafting pays, predicting the longest accepted blocks on grounded tasks, where the target's next-token entropy is lowest. Our code is available at https://github.com/js-lee-AI/GLANCE.

Comments21 pages, 8 figures, 16 tables. Code: https://github.com/js-lee-AI/GLANCE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑