arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlashDiff:用于扩散模型服务的高效区域执行与调度

FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving

Yaqi Qiao, Ping He, Songrun Xie, Ayush Barik, Chensong Zhang, Zhengzhong Tu, Fan Lai

arXiv 2607.12121首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Vanderbilt University; HKUST; NVIDIA; Texas A&M University(伊利诺伊大学厄巴纳-香槟分校; 范德比大学; 香港科技大学; 英伟达; 德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对扩散模型服务效率低的问题,提出FlashDiff系统,通过自适应区域执行和调度,利用扩散细化特性及三种机制,有效降低端到端服务延迟,提高吞吐量。

AI 中文摘要

扩散模型已成为现代图像、视频和音频生成的核心支柱,但其高效服务仍是挑战。与自回归解码不同,扩散推理在多个去噪步骤中反复更新高维空间或时间潜在变量。现有多GPU并行化方法存在问题。本文提出FlashDiff,通过自适应区域执行和调度提高推理效率。基于扩散细化在潜在区域或去噪步骤中不均匀的观察,FlashDiff利用这些特性选择性执行需进一步细化的区域,并在并发服务请求中重新分配计算空闲时间。它由三种机制组成,包括使用早期注意力信号分解潜在表示、用轻量级运行时控制器估计区域活动、应用亲和感知在线调度器。在实际图像、视频和音频工作负载中,FlashDiff将端到端服务延迟降低30 - 97%,吞吐量提高1.2 - 2.2倍。

英文摘要

Diffusion models have become the central backbone for modern image, video, and audio generation, but their efficient service remains a challenge. Unlike autoregressive decoding, diffusion inference repeatedly updates high-dimensional spatial or temporal latents over many denoising steps. This all-region execution pattern makes generation latency high and limits serving throughput. Existing multi-GPU parallelization methods can reduce per-step computation, but often introduce substantial activation exchange overhead, causing communication to offset or even outweigh the benefits of parallel execution. This paper presents FlashDiff, a diffusion serving system that improves inference efficiency through adaptive regional execution and scheduling. FlashDiff is based on the observation that diffusion refinement is not uniform across latent regions or denoising steps: different regions often stabilize at different rates, while neighboring steps exhibit strong temporal correlation. FlashDiff leverages these properties to selectively execute only regions that require further refinement and to reallocate the resulting compute slack across concurrent serving requests. FlashDiff consists of three mechanisms. First, it decomposes the latent representation into coherent execution regions using early-stage attention signals, preserving semantic structure while exposing fine-grained parallelism. Second, it uses a lightweight runtime controller to estimate region activity and bypass low-impact updates when further refinement is unlikely to affect output quality. Third, it applies an affinity-aware online scheduler that co-locates dependent regions, balances residual load across GPUs, and reuses reclaimed compute capacity to improve serving efficiency. Across real-world image, video, and audio workloads, FlashDiff reduces end-to-end serving latency by 30-97% and improves throughput by 1.2-2.2x.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑