发表机构
Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SpecFold通过折叠多分支投机验证中的计算冗余,结合算法与系统协同设计,在扩散语言模型上实现高达1.99倍的解码加速,同时保持任务性能。
AI 中文摘要
扩散大语言模型(DLLMs)通过迭代块去噪生成文本,而多分支投机解码通过在单次前向传播中同时验证主分支和多个草稿分支来加速这一过程。虽然先前的DLLM加速方法主要利用去噪步骤间的时间冗余,但我们识别出每个投机验证步骤内一个互补的冗余轴:多分支计算冗余。在投机验证期间,草稿分支从父分支继承大部分令牌,同时仅解掩码少量额外位置,导致跨分支的隐藏状态大部分高度相似。我们提出SpecFold,一种算法-系统协同设计,利用这种多分支冗余来降低多分支投机验证的成本。在算法层面,SpecFold执行令牌级残差门控,并通过折叠注意力和FFN选择性重用父计算,同时保留残差隐藏状态。在系统层面,Triton内核实现通过高效稀疏多分支执行,将这种细粒度重用转化为端到端吞吐量提升。SpecFold与时间缓存正交,并与现有DLLM投机策略兼容。在两种DLLM家族、五个模型和五个标准基准上,SpecFold相对于Spiffy实现高达1.64倍的吞吐量提升,相对于普通解码实现高达1.99倍的提升,同时保持相当的任务性能。
英文摘要
Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch computational redundancy. During speculative verification, draft branches inherit most tokens from their parents while unmasking a small set of additional positions, causing large portions of hidden states to remain highly similar across branches. We propose SpecFold, an algorithm-system co-design that exploits this multi-branch redundancy to reduce the cost of multi-branch speculative verification. Algorithmically, SpecFold performs token-level residual gating and selectively reuses parent computation through folded attention and FFN while preserving residual hidden states. Systemically, a Triton kernel implementation translates this fine-grained reuse into end-to-end throughput gains through efficient sparse multi-branch execution. SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies. Across two DLLM families, five models, and five standard benchmarks, SpecFold achieves up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding, while maintaining comparable task performance.