arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12829cs.LGcs.AIcs.CL

加速掩码扩散大语言模型:高效推理技术综述

Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques

Daehoon Gwak, Minhyung Lee, Junwoo Park, Jaegul Choo

首次发表
浏览论文内容

中文总结 AI 辅助

综述介绍用于扩散大语言模型的统一延迟分解框架,以理清算法、架构和系统因素对推理速度的影响,将加速技术分类,提供可重复基准测试指导方针并强调实现并行生成潜力的挑战。

中文摘要 AI 辅助

扩散大语言模型(dLLMs)在并行生成方面相对于标准自回归模型具有理论优势。然而,仅并行生成并不能保证实际加速。实现这种效率需要专门的推理机制,如扩散感知缓存和重用。随着推理效率成为实际部署的先决条件,近期研究积极探索跨算法、架构和系统的加速技术。但由于现有基准测试中算法、架构和系统级因素之间复杂的权衡导致端到端延迟难以进行严格比较。本综述引入统一延迟分解框架来理清这些因素并分析其对实际部署中推理速度的影响。在此框架指导下,将加速技术沿算法创新、架构与系统优化以及推理时间缩放三个轴进行分类。最后提供可重复基准测试的指导方针并强调实现并行生成全部潜力的开放挑战。

英文摘要

Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as diffusion-aware caching and reuse. Consequently, as inference efficiency becomes a prerequisite for practical deployment, recent research has actively explored acceleration techniques across algorithms, architectures, and systems. However, rigorous comparisons remain difficult, as end-to-end latency stems from intricate trade-offs between algorithmic, architectural, and system-level factors that are often conflated in existing benchmarks. In this survey, we introduce a unified latency decomposition framework for dLLMs to disentangle these factors and analyze their impact on inference speed in real deployments. Guided by this framework, we categorize acceleration techniques along three axes covering algorithmic innovations, architectural and system optimizations, and inference-time scaling. Finally, we provide guidelines for reproducible benchmarking and highlight open challenges for realizing the full potential of parallel generation.

发表机构

  • KAIST AI(韩国科学技术院人工智能研究所)
  • Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑