arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考少步扩散Transformer中的缓存目标:求解器感知的目标选择

Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection

Shuo Yang, Lihao Fang, Yi Zhang, Haixiang Wang, Xincheng Ye, Shufan Chen, Jipeng Guo, Youqing Wang

arXiv 2610.03577首次发表:更新:

AI 中文总结

针对少步扩散Transformer缓存目标选择问题,提出AutoTarget方法,通过测量候选张量重用误差选择最优缓存,减少计算开销并保持生成质量。

AI 中文摘要

扩散Transformer(DiTs)能够生成高质量的图像和视频,但生成每个样本需要多次昂贵的DiT前向传递。加速DiT采样的两种常见方法是步蒸馏(减少采样步数)和缓存(通过重用先前步骤计算的张量来跳过某些DiT评估)。大多数缓存方法预先决定重用哪个张量。蒸馏后,相邻采样步之间的距离更远。跨越这一更大间隔重用张量会引入更多误差,因此选择缓存内容变得尤为重要。为此,我们引入了AutoTarget,一种为给定模型、求解器和重用调度选择缓存张量的方法。AutoTarget使用一组无缓存重用的小规模运行来测量重用每个候选张量引起的误差,然后选择误差最低的候选。我们还分析了重用步骤中的误差如何影响最终样本。对于Euler采样,我们识别出能产生相同轨迹的缓存目标,并解释了为何存储的求解器更新可能无法做到这一点。在蒸馏图像和视频DiT上的实验表明,最佳缓存目标随模型、图像分辨率和求解器而变化。AutoTarget减少了DiT评估次数和保留的缓存存储量。生成质量仍接近对应的无缓存运行。在测试的PixArt-LCM和FLUX.1-schnell设置中,其校准排名与保留缓存运行的排名一致。为帮助他人复现该方法,我们在GitHub上提供了其核心实现,网址为https://this https URL。

英文摘要

Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluations by reusing a tensor computed at an earlier step. Most caching methods decide in advance which tensor to reuse. After distillation, adjacent sampling steps are farther apart. Reusing a tensor across this larger gap introduces more error, so choosing what to cache becomes especially important. We therefore introduce AutoTarget, a method that chooses the cached tensor for a given model, solver, and reuse schedule. AutoTarget uses a small set of runs without cache reuse to measure the error caused by reusing each candidate tensor, then selects the candidate with the lowest error. We also analyze how an error at one reuse step affects the final sample. For Euler sampling, we identify cache targets that produce the same trajectory and show why a stored solver update may not. Experiments on distilled image and video DiTs show that the best cache target changes with the model, image resolution, and solver. AutoTarget reduces DiT evaluations and retained cache storage. Generation quality remains close to the corresponding uncached run. On the tested PixArt-LCM and FLUX.1-schnell settings, its calibration ranking matches the ranking from held-out cached runs. To help others reproduce the method, we provide its core implementation on GitHub at https://github.com/wali1024-offical/AutoTarget.

Comments20 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑