发表机构
National Taiwan University; University of Southern California; Artificial Intelligence Center of Research Excellence, National Taiwan University(国立台湾大学; 南加州大学; 国立台湾大学卓越人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散模型并行采样中重叠时间步导致的冗余计算,提出与算法无关的输出缓存机制ParaAnya,重用缓存结果并仅分派未命中时间步,在Stable Diffusion v1.5上实现最高5.62倍加速且保持质量。
AI 中文摘要
扩散模型在生成任务中取得了显著成功,但其固有的顺序采样过程引入了严重的计算瓶颈。最近的并行时间(PinT)求解器试图通过在时间步的滑动窗口上并行化生成来缓解这一问题,仅当逐步变化稳定时才推进窗口。然而,这种重叠窗口机制迫使网络重复评估相同的时间步。当迭代之间的输入变化最小时,这些冗余评估导致显著的计算浪费。为了解决这一低效问题,我们提出了ParaAnya,一种与并行采样算法无关的输出缓存机制,能够减少函数评估次数(NFE)。ParaAnya缓存扩散模型的输入-输出对,并在重叠时间步重用缓存的输出。通过仅将缓存未命中的时间步分派给GPU工作线程,我们的方法消除了冗余计算,同时保留了底层算法更新规则的结构。我们将ParaAnya集成到四种代表性的并行采样算法中,并在Stable Diffusion v1.5上评估其性能。在八块GPU上使用DDIM评估的四种并行采样器中,ParaAnya相对于未缓存版本提供了1.30至2.43倍的加速,并将NFE减少了高达70.1%,相比单GPU串行采样达到了高达5.62倍的加速,同时保持了相当的CLIP分数。
英文摘要
Diffusion models have achieved remarkable success in generative tasks, but their inherently sequential sampling process introduces a severe computational bottleneck. Recent Parallel-in-Time (PinT) solvers attempt to mitigate this by parallelizing generation across a sliding window of timesteps, advancing the window only when step-wise changes stabilize. However, this overlapping window mechanism forces the network to repeatedly evaluate the same timesteps. When the input variations between iterations are minimal, these redundant evaluations lead to significant computational waste. To address this inefficiency, we propose ParaAnya, an output cache mechanism agnostic to the parallel sampling algorithm that can reduce the number of function evaluations (NFE). ParaAnya caches input-output pairs of diffusion models and reuses the cached output at overlapping timesteps. By dispatching only cache-miss timesteps to GPU workers, our approach eliminates redundant computation while preserving the structure of the underlying algorithms' update rules. We integrate ParaAnya into four representative parallel sampling algorithms and evaluate its performance on Stable Diffusion v1.5. Across four parallel samplers evaluated with DDIM on eight GPUs, ParaAnya provides $1.30$--$2.43\times$ speedups over their uncached counterparts and reduces NFE by up to 70.1\%, reaching up to a $5.62\times$ speedup over single-GPU serial sampling while maintaining comparable CLIP scores.
Comments5 pages