发表机构
Data Science & Artificial Intelligence Research Institute, China Unicom; Unicom Data Intelligence, China Unicom(中国联通数据科学与人工智能研究院; 中国联通数据智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过调查多模态模型并开展跨架构实证研究,探究多模态推测解码是否适用于基于扩散的并行草稿生成,提出统一分类法并分析现有方法的局限性与未来方向。
AI 中文摘要
推测解码通过让轻量级草稿模型提出未来 token,同时由目标模型并行验证,从而加速自回归生成,其无损保证推动了一系列将草稿模型自身向并行生成方向发展的研究。最新范式是块并行生成草稿,包括 DFlash、DSpark 等基于扩散的方法,在日常聊天任务上实现了最高 3.6 倍的加速。尽管这一转变在纯文本大语言模型中已得到充分研究,但其在多模态模型中的适用性仍是开放问题。现有多模态推测解码工作聚焦于输入压缩、适配器对齐、候选覆盖或模态特定验证,而块并行生成草稿仍未被充分探索。为弥合这一差距,本文结合以模态为中心的调查与跨架构实证研究,探究多模态推测解码是否适用于基于扩散的并行草稿生成。在调查中,我们系统分析了涵盖视觉-语言、视频-语言、音频、视觉-语言-动作(VLA)架构的多模态模型,从草稿并行性与跨模态信息交互双视角展开。我们提出统一分类法,将草稿模型侧并行性与树构建、验证策略等正交设计选择分离。此外,我们在标准化多模态基准(包括 OCR、VQA、视觉推理、图像字幕)上,对不同并行度下的现有方法进行全面实证比较。最后,我们总结当前方法的局限性,讨论开放挑战,并为这一快速发展领域勾勒有前景的未来方向。
英文摘要
Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work that pushes the drafter itself toward parallel generation. The most recent paradigm is block-parallel generative drafting, including diffusion-based methods such as DFlash and DSpark, achieving up to 3.6x speedup on common daily chatting tasks. While this transition is well studied in text-only LLMs, its applicability to multimodal models remains an open question. Existing multimodal speculative decoding efforts focus on input compression, adapter alignment, candidate coverage, or modality-specific verification; however, block-parallel generative drafting remains largely unexplored. To bridge this gap, this paper combines a modality-centered survey with a cross-architecture empirical study to ask: Is multimodal speculative decoding ready for diffusion-based parallel drafting? In this survey, we systematically analyze a wide spectrum of multimodal models, spanning Vision-Language, Video-Language, Audio, and Vision-Language-Action (VLA) architectures, from the dual perspectives of drafting parallelism and cross-modal information interaction. We introduce a unified taxonomy that isolates drafter-side parallelism from orthogonal design choices such as tree construction and verification strategies. Furthermore, we provide a comprehensive empirical comparison of existing methods under varying degrees of parallelism across standardized multimodal benchmarks, including OCR, VQA, visual reasoning, and image captioning. Finally, we summarize the limitations of current approaches, discuss open challenges, and outline promising future directions for this rapidly evolving field.