arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Xema:通过细粒度内存管理和自动配置实现高效扩散服务

Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration

Xueze Kang, Guangyu Xiang, Suyi Li, Yuxin Wang, Shaohuai Shi, Lin Zhang, Xiaowen Chu

arXiv 2607.11136首次发表:更新:

AI 中文总结

研究针对扩散模型服务受GPU内存限制问题,提出Xema系统,通过细粒度内存管理和自动配置,利用可预测张量生命周期优化内存,引入离线规划器联合选择参数,实现高效扩散服务,相比现有配置提升SLO达成率并降低规划成本。

AI 中文摘要

扩散模型越来越多地作为生产视觉生成服务进行部署,在服务高分辨率图像和长视频生成时,GPU内存常常成为限制因素。诸如权重卸载、分片和VAE切片等常用内存节省技术往往不实用,因为会带来显著性能开销。本文提出Xema,一个内存高效的扩散服务系统,利用可预测张量生命周期进行跟踪引导的内存优化。对于每个请求模板,Xema导出离线内存跟踪以识别短内存压力区间,并仅在这些区间内按所需量应用内存缓解以符合目标GPU预算。Xema还为具有可预测生命周期的张量构建静态内存布局,减少碎片化导致的预留内存,并使离线内存推理在运行时可靠。在此内存优化层之上,Xema引入离线规划器,在GPU内存和SLO约束下联合选择并行性、并发性和内存控制。所选计划存储在计划表中并由在线服务运行时直接使用。我们在生产扩散管道上实现了Xema,并使用Flux.2、CogVideoX - 5B和LTX - 2进行评估。与现有服务配置相比,Xema将SLO达成率提高了3.7倍,并将规划成本从6.3小时降至197秒。

英文摘要

Diffusion models are increasingly deployed as production visual-generation services, where serving high-resolution image and long video generation is often limited by GPU memory. Popular memory-saving techniques such as weight offloading, sharding, and VAE slicing are often not practical because they tend to introduce significant performance overhead. In this paper, we present Xema, a memory-efficient diffusion serving system that exploits predictable tensor lifetimes for trace-guided memory optimization. For each request template, Xema derives an offline memory trace to identify short memory-pressure intervals and applies memory mitigation only within these intervals and only by the amount needed to fit the target GPU budget. Xema further constructs a static memory layout for tensors with predictable lifetimes, reducing fragmentation-induced reserved memory and making offline memory reasoning reliable at runtime. Built on this memory optimization layer, Xema introduces an offline planner that jointly selects parallelism, concurrency, and memory control under GPU memory and SLO constraints. The selected plan is stored in a plan table and directly used by the online serving runtime. We implement Xema on production diffusion pipelines and evaluate it with Flux.2, CogVideoX-5B, and LTX-2. Compared with existing serving configurations, Xema improves SLO attainment by up to 3.7x and reduces planning cost from 6.3 hours to 197 seconds compared with grid search.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑