arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DynaTE:通过动态令牌执行加速扩散大语言模型

DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution

Minghan Jiang, Jiayi Wang, Shuaiting Li, Haibin Shen, Kejie Huang

arXiv 2610.11284首次发表:更新:

发表机构

College of Information Science & Electronic Engineering, Zhejiang University(浙江大学信息与电子工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出软硬件协同设计架构DynaTE,通过动态调整加速器执行适配dLLM解码,实现2.05-2.78倍加速及更高能效,优于现有dLLM加速器与Jetson AGX Orin。

AI 中文摘要

基于扩散的大语言模型(dLLM)近期成为自回归(AR)大语言模型的有前景替代方案,它支持双向并行优化,缓解了自回归生成的顺序解码瓶颈。然而,其并行迭代优化与为顺序解码优化的AR加速器不匹配,且离散令牌生成与为连续去噪设计的DiT加速器存在差异。近期的dLLM加速器已探索针对特定工作负载的优化,以降低词汇处理开销和去噪迭代间的冗余计算,但这些方法在并行执行中仍保留所有令牌,尽管不同令牌的优化效用和执行需求存在差异。本文提出DynaTE,一种软硬件协同设计架构,可在dLLM解码过程中动态调整加速器执行以适配不断变化的令牌状态。DynaTE首先通过跳过低效用令牌计算实现自适应令牌执行,同时维度可重构PE阵列在不同活跃令牌模式下保持高利用率;其次,DynaTE通过FLDD利用动态令牌依赖关系,在当前迭代中优化少量局部依赖令牌,减少整体去噪迭代次数,而Merge-Split-Merge数据流可隐藏由此产生的串行开销;第三,流式词汇引擎交错来自LM头的多个令牌流,以适配选择性令牌计算和不均衡词汇选择需求导致的不规则输出变化。在两个代表性dLLM上评估显示,与最先进的dLLM加速器相比,DynaTE实现2.05至2.78倍的加速和2.99至3.93倍的能效提升,与Jetson AGX Orin相比则实现2.55倍加速和6.07倍能效提升。

英文摘要

Diffusion-based LLMs (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs by enabling bidirectional parallel refinement, alleviating the sequential decoding bottleneck of AR generation. However, their parallel iterative refinement mismatches AR accelerators optimized for sequential decoding and their discrete token generation differs from DiT accelerators designed for continuous denoising. Recent dLLM accelerators have explored workload-specific optimizations to reduce vocabulary processing overhead and redundant computation across denoising iterations. However, these approaches retain all tokens in parallel execution, despite varying token refinement utility and execution requirements. This paper presents DynaTE, a hardware--software co-design architecture that dynamically adapts accelerator execution to evolving token states during dLLM decoding. DynaTE first enables adaptive token execution by skipping low-utility token computation, while a dimension-reconfigurable PE array maintains high utilization under varying active-token patterns. Second, DynaTE exploits dynamic token dependencies through FLDD to refine a small number of locally dependent tokens within the current iteration, reducing the overall number of denoising iterations, while a Merge--Split--Merge dataflow hides the resulting serial overhead. Third, a streaming vocabulary engine interleaves multiple token streams from the LM head to accommodate irregular output variations caused by selective token computation and uneven vocabulary-selection demands. Evaluated on two representative dLLMs, DynaTE achieves 2.05--2.78$\times$ speedup and 2.99--3.93$\times$ higher energy efficiency over state-of-the-art dLLM accelerators, while delivering 2.55$\times$ speedup and 6.07$\times$ higher energy efficiency over Jetson AGX Orin.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑