发表机构
University of California, Los Angeles; Samsung Research America(加州大学洛杉矶分校; 三星美国研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HAWK通过表示相似性选择目标层、压缩视觉隐藏状态并建模目标预测变化,提升多模态推测解码的接受长度与加速比。
AI 中文摘要
推测解码已为大型语言模型(LLMs)实现了显著的、无损的加速,但对于大型视觉-语言模型(LVLMs)效果仍然不佳,因为轻量级草稿模型难以利用丰富的多模态信息。第二个限制是,标准蒸馏仅沿原始训练轨迹监督草稿模型,而未建模在草稿模型自身提出建议后目标预测如何变化。随着草稿偏离该轨迹,草稿模型与目标模型的分歧可能越来越大,从而降低后续步骤的接受率。我们提出HAWK来解决这两个限制。HAWK使用表示相似性来选择信息量大的目标层,并学习如何组合它们的隐藏状态。对于视觉信息,它直接向草稿模型提供来自目标模型的压缩视觉隐藏状态,而非原始视觉标记,使得浅层草稿模型更容易利用视觉信息。HAWK还训练草稿模型捕捉其自身建议后目标预测的变化,从而提高多步草稿过程中与目标模型的一致性。在SmolVLM-256M上,跨十个多模态基准,在贪婪解码下,HAWK将平均接受长度从3.32提升至4.08,加速比从2.19倍提升至2.60倍(相对于EAGLE-3);在采样下,分别从2.89提升至3.41,以及从1.92倍提升至2.19倍。
英文摘要
Speculative decoding has achieved substantial lossless speedups for LLMs, but remains less effective for large vision-language models (LVLMs), where lightweight drafters struggle to use rich multimodal information. A second limitation is that standard distillation supervises the drafter only along the original training trajectory, without modeling how target predictions shift after the drafter's own proposals. As drafting moves away from this trajectory, the drafter can increasingly disagree with the target, reducing acceptance in later steps. We propose HAWK to address both limitations. HAWK uses representation similarity to select informative target layers and learns how to combine their hidden states. For visual information, it directly provides the drafter with compressed visual hidden states from the target model instead of raw visual tokens, making the visual information easier for a shallow drafter to use. HAWK also trains the drafter to capture how target predictions change after its own proposals, improving its agreement with the target during multi-step drafting. On SmolVLM-256M across ten multimodal benchmarks, HAWK raises average acceptance length from 3.32 to 4.08 and speedup from 2.19x to 2.60x over EAGLE-3 under greedy decoding, and from 2.89 to 3.41 and 1.92x to 2.19x under sampling.