发表机构
College of Electronics and Information Engineering, Shenzhen University; School of Artificial Intelligence, Shenzhen University; College of Computer Science and Software Engineering, Shenzhen University(深圳大学电子与信息工程学院; 深圳大学人工智能学院; 深圳大学计算机与软件学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出无训练加速方法BaryCache,采用重心外推器缓解扩散Transformer采样的振荡伪影,在图像与视频生成中实现最高3.30倍采样加速,兼顾内存与感知质量。
AI 中文摘要
扩散Transformer(DiT)可实现高保真图像与视频生成,但其迭代采样因每步去噪需大量矩阵运算而成本高昂。现有基于缓存的加速方法虽能减少冗余计算,却因存储中间状态增加了显存(VRAM)占用,直接限制推理批次大小。本研究提出一种无训练加速方法,采用重心外推器(Barycentric Extrapolator)对DiT采样进行逐步预测;通过利用重心外推,该预测器数值稳定,可缓解前向预测中类似龙格现象的振荡伪影。在图像与视频生成的大量实验中,本方法在内存使用与感知质量间取得良好权衡,相比基线DiT推理实现了最高3.30倍的端到端采样加速。
英文摘要
Diffusion Transformers achieve high-fidelity image and video generation, but their iterative sampling remains expensive, for each denoising step requires large matrix operations. Existing cache-based acceleration reduces redundant computation yet increases the VRAM footprint by storing intermediate states, which can directly constrain inference batch size. In this work, we propose a training-free acceleration method that performs stepwise forecasting for DiT sampling using a Barycentric Extrapolator. By leveraging barycentric extrapolation, our predictor is numerically stable and alleviates oscillatory artifacts analogous to the Runge phenomenon during forward forecasting. Across extensive experiments on both image and video generation, our approach provides a favorable trade-off between memory usage and perceptual quality, while delivering up to 3.30x end-to-end sampling speedup compared with baseline DiT inference.
CommentsAccepted by ICITES 2026