arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DSTAR:通过减少空间和时间冗余加速扩散变换器

DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction

Chi Zhang, Jieru Zhao, Yu Feng, Chen Zhang, Quan Chen, Minyi Guo

arXiv 2607.15846首次发表:更新:

AI 中文总结

研究针对扩散变换器多迭代推理低效耗能问题,提出软硬件协同设计框架DSTAR。算法上用细粒度混合精度量化和稀疏注意力重用机制,架构上设计专用硬件加速器,经评估在多个典型DiTs上实现显著延迟加速和节能,且不降低精度。

AI 中文摘要

扩散变换器(DiTs)已广泛应用于图像合成、视频生成和内容编辑等任务。但其多迭代推理过程导致性能低效和高能耗。现有加速方法主要关注减少相邻时间步之间的时间冗余,却常忽略DiTs的特定特征。我们提出DSTAR,一个软硬件协同设计框架,通过减少空间和时间冗余加速DiT推理。算法层面,引入细粒度混合精度量化方法,增加低比特计算比例,还采用稀疏注意力重用机制减少注意力层冗余计算。架构上设计专用硬件加速器。在七个典型DiTs上评估表明,与NVIDIA A100 GPU相比,DSTAR延迟加速达7.33倍、节能41.89倍,与SOTA加速器相比,延迟加速达2.54倍、节能3.68倍,且无精度下降。

英文摘要

Diffusion Transformers (DiTs) have been widely used in many tasks, including image synthesis, video generation, and content editing. However, their multi-iteration inference process leads to performance inefficiency and high energy consumption. Existing acceleration methods primarily focus on reducing temporal redundancy between adjacent timesteps, but often overlook the specific features of DiTs. As a result, these approaches either suffer from great accuracy degradation or fail to achieve high efficiency. We present DSTAR, a software-hardware co-design framework that accelerates DiT inference by reducing spatial and temporal redundancy. At the algorithmic level, DSTAR introduces a fine-grained mixed-precision quantization method for differential activations in linear operations, significantly increasing the proportion of low-bit computations. Additionally, DSTAR incorporates a sparse attention reuse mechanism to minimize redundant computation in attention layers. For architectural support, we design a specialized hardware accelerator which achieves high efficiency in both latency and energy consumption. Evaluation on seven typical DiTs demonstrates that DSTAR achieves up to 7.33x latency speedup and 41.89x energy savings compared to an NVIDIA A100 GPU, and achieves up to 2.54x latency speedup and 3.68x energy savings compared to SOTA accelerators, without accuracy degradation.

CommentsAccepted to the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑