arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25380cs.AR

APT:基于注意力概率引导剪枝与量化的扩散Transformer加速方案

APT: Accelerating Diffusion Transformers via Attention Probability-Guided Pruning and Quantization

Sungyeob Yoo, Seeyeon Kim, Joonyong Park, Seunghee Han, Joo-Young Kim

AI总结:

本文提出软硬件协同设计的APT加速器,通过APDT与TAFA优化DiTs的自注意力计算,在多款SOTA DiT模型上实现显著加速与能效提升。

AI中文摘要:

生成式AI的近期进展大幅提升了高分辨率图像与视频生成的需求,使扩散模型成为核心技术。其中,扩散Transformer(DiTs)因可扩展性与输出质量成为当前最优(SOTA)模型,但DiTs中的自注意力会产生显著计算开销,且延迟随输出分辨率四次方增长而过长。现有工作虽尝试用稀疏性与量化技术降低成本,但无法有效减少高分辨率DiTs的计算开销。本文提出APT,一种面向高分辨率DiTs的软硬件协同设计加速器。APT以注意力概率为统一重要性指标,通过细粒度剪枝与自适应精度缩放联合优化计算:算法层面提出注意力概率引导的自适应双阈值(APDT),利用双阈值动态执行元素选择与精度分配;为确保与内存高效的FlashAttention兼容,引入时间步感知FlashAttention(TAFA),通过利用时间相似度预测各时间步的注意力概率。架构层面协同设计专用加速器,高效支持不规则稀疏性与双精度执行,具备动态掩码管理、地址转换、双精度计算单元及基于分块的数据流。最后在PixArt-α、Stable Diffusion 3、FLUX等SOTA DiT模型上评估APT,结果显示其相比NVIDIA A100实现最高8.16倍加速与14.98倍能效提升,相比SOTA扩散模型加速器EXION实现最高3.01倍加速与2.04倍能效提升。

英文摘要:

Recent advances in generative AI have significantly increased the demand for high-resolution image and video generation, positioning diffusion models as a core technology. Among them, Diffusion Transformers (DiTs) have emerged as the state-of-the-art (SOTA) models due to their scalability and output quality. However, self-attention in DiTs incurs significant computational overhead, leading to excessively long latency as the complexity grows with the fourth power of the output resolution. While prior works have attempted to mitigate this cost using sparsity and quantization techniques, they fall short of effectively reducing the computational cost in high-resolution DiTs. In this paper, we present APT, a software-hardware co-designed accelerator for high-resolution DiTs. APT leverages attention probabilities as a unified importance metric to jointly optimize computation through fine-grained pruning and adaptive precision scaling. At the algorithm level, we propose Attention Probability-guided Adaptive Dual Thresholding (APDT), which dynamically performs element selection and precision assignment using dual thresholds. To ensure compatibility with memory-efficient FlashAttention, we introduce Timestep-Aware FlashAttention (TAFA), which predicts attention probabilities across timesteps by exploiting temporal similarity. At the architecture level, we co-design a specialized accelerator that efficiently supports irregular sparsity and dual-precision execution, featuring dynamic mask management, address translation, dual-precision compute units, and a tile-based dataflow. Finally, we evaluate APT on SOTA DiT models, including PixArt-$α$, Stable Diffusion 3, and FLUX. APT achieves up to 8.16$\times$ speedup and 14.98$\times$ higher energy efficiency over NVIDIA A100, and up to 3.01$\times$ speedup and 2.04$\times$ higher energy efficiency over EXION, a SOTA diffusion model accelerator.

补充信息

↑