arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05981cs.CVcs.LG

加速扩散Transformer:基于高斯过程修正的特征缓存

Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache

发表机构上海交通大学 · 电子科技大学 · 山东大学
另 3 家 · 查看机构详情
  • Shanghai Jiao Tong University(上海交通大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • Shandong University(山东大学)
  • Xiamen University(厦门大学)
  • Terminal Intelligent Computing Division, Alibaba Cloud(阿里云终端智能计算部)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhirong Shen, Rui Huang, Chang Zou, Shikang Zheng, Jiacheng Liu, Peiliang Cai, Zhengyi Shi, Yaosong Du, Liang Feng, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Linfeng Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散Transformer加速中预测偏差问题,提出基于高斯过程回归的即插即用修正框架,通过不确定性自适应策略降低计算负载并提升生成质量。

中文摘要 AI 辅助

扩散Transformer已成为生成式AI中的主导范式,但其高昂的计算成本严重阻碍了实时应用。基于预测的特征缓存被广泛用于加速扩散Transformer,然而,随着步数增加,其预测与参考全计算轨迹之间的偏差逐渐增大。一个直观的想法是使用在线回归模型动态修正这一偏差,但这面临加速过程中标签数据不可用的问题。本文提出一个统计观察:使用缓存方法的全计算步特征与参考全计算轨迹之间的残差局部呈现零均值高斯分布。通过将全计算步的特征视为参考特征的有噪声观测,数据获取问题得以解决。基于此观察,本文提出一个即插即用的GP-Refiner修正框架。该方法利用高斯过程回归进行修正,并利用GPR的性质,引入一种不确定性自适应计算策略,通过实时监测后验方差来触发必要的全计算校准。实验表明,与多种最先进方法结合时,在不同模型上均有显著改进。将所提框架与TaylorSeer集成,计算负载降低19.3%,同时PSNR提高0.9 dB,LPIPS从0.46降至0.29。代码可在该https URL获取。

英文摘要

Diffusion Transformers have become the dominant paradigm in generative AI, but their high computational costs severely hinder real-time applications. Prediction-based feature caching is widely used to accelerate diffusion transformers; however, as the number of steps increases, the deviation between its predictions and the reference full-compute trajectory gradually grows. An intuitive idea is to use an online regression model to dynamically correct this deviation, but it faces the issue of label data being unavailable during the acceleration process. This paper presents a statistical observation that the residuals between the features of full computation steps using caching methods and reference full-compute trajectory locally exhibit a zero-mean Gaussian distribution. By treating the features of full computation steps as noisy observations of reference features, the data acquisition problem is resolved. Based on this observation, a plug-and-play GP-Refiner correction framework is proposed. This method utilizes Gaussian Process Regression for correction and, leveraging the properties of GPR, introduces an uncertainty-adaptive computation strategy that triggers necessary full-computation calibration by monitoring the posterior variance in real time. Experiments demonstrate significant improvements across different models when combined with various state-of-the-art methods. Integrating the proposed framework with TaylorSeer reduces the computational load by 19.3% while improving PSNR by 0.9 dB and reducing LPIPS from 0.46 to 0.29. Code is available in https://github.com/Aredstone/GP-Refiner.

补充信息

相关深度报道

↑