发表机构
Adobe; NVIDIA(Adobe公司; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对消费级GPU上的扩散模型推理,提出嵌入翻译器、扫描配方和交互式编辑器,实现亚秒级首令牌时间,兼顾性能、质量与模型占用。
AI 中文摘要
设备端推理正在蓬勃发展,但这一势头几乎全部集中在语言模型上。扩散流水线内存占用大、对延迟敏感,并且需要协调嵌入器、变换器、解码器,以及通常进一步的后处理,这些后处理并不像大语言模型推理循环那样标准化。我们在性能、质量和模型占用空间之间进行权衡,以尽可能覆盖现实世界中的众多客户端设备。我们做出了三项贡献:一个嵌入翻译器,将小型文本编码器映射到大型编码器空间,以削减权重和延迟;一个可复现的扫描配方,用于在扩散流水线中导航速度/质量/内存三角;以及一个交互式设备端图像生成编辑器,在近期GPU上实现亚秒级TTFI。
英文摘要
On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a decoder, and often further postprocessing that is not as standardized as LLM inference loops are. We navigate the trade-off between performance, quality, and model footprint to reach as many client devices in the wild as possible. We make three contributions: an embedding translator that maps a small text encoder into a large encoder space to cut weight and latency; a reproducible sweep recipe for navigating the speed/quality/memory triangle in diffusion pipelines; and an interactive on-device image generation editor achieving sub-second TTFI on recent GPUs.