发表机构
Harbin Institute of Technology, Shenzhen; Pengcheng Laboratory(哈尔滨工业大学(深圳); 鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对SJD加速自回归图像生成中的令牌歧义问题,本文提出SJD-SV方法,通过语义感知令牌子序列级验证替代逐令牌验证,即插即用,显著提升现有SJD方法的性能。
AI 中文摘要
推测雅可比解码(SJD)是加速自回归图像生成的重要方法。尽管SJD已展现出优越性能,但近期研究指出,其在令牌验证过程中常遭受令牌歧义问题,且其原因难以得到充分解释。为探究这一原因,本文对视觉令牌进行了可视化分析,发现与文本令牌不同,视觉令牌通常对应一些局部、微小且不清晰的视觉细节,这意味着仅使用单个令牌难以准确表达某种语义,从而引发令牌歧义问题。为此,我们提出了一种新颖的带语义验证的推测雅可比解码方法(称为SJD-SV),用于加速自回归图像生成。其核心思想是利用令牌之间强大的修正特性来识别语义感知的令牌子序列,然后不再进行逐令牌验证,而是转向在语义感知的令牌子序列层面进行验证,以加速图像生成。特别地,我们的方法是即插即用的,可直接集成到现有SJD及其变体中。在多个数据集上的大量实验表明,现有SJD方法在集成我们的SJD-SV方法后,性能均获得显著提升。
英文摘要
Speculative Jacobi Decoding (SJD) is an important approach for accelerating autoregressive image generation. Although SJD has shown superior performance, recent studies point out that it usually suffers from a token ambiguity issue during token verification but its reason can not be well explained. To figure out this reason, in this paper, we conduct a visualization analysis on vision token and find that different from text tokens, vision tokens generally corresponds to some local, small, and unclear vision details, which means only using single token is difficult to accurately express a certain semantic, thereby causing token ambiguity issue. To this end, we propose a novel Speculative Jacobi Decoding with Semantics Verification (called SJD-SV), for accelerating autoregressive image generation. The key idea is that leveraging the strong correction characters between tokens to recognize semantic-aware token subsequence and then instead of perform token-by-token verification, turning to perform verification on semantic-aware token subsequence level for accelerating image generation. In particular, our method is plug-in, which can be directly integrated into existing SJD and its variants. Extensive experiments on various datasets show that existing SJD methods achieve significant performance improvement after integrating our SJD-SV method.
CommentsAccepted at the 43rd International Conference on Machine Learning (ICML 2026)
Journal refProceedings of the 43rd International Conference on Machine Learning, PMLR 306, 2026