arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重量终结——消费级GPU上的交互式扩散

The Weight Is Over - Interactive Diffusion on Consumer GPUs

Frieder Ganz, Maximilian Müller

arXiv 2609.21849首次发表:更新:

发表机构

Adobe; NVIDIA(Adobe公司; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对消费级GPU上的扩散模型推理,提出嵌入翻译器、扫描配方和交互式编辑器,实现亚秒级首令牌时间,兼顾性能、质量与模型占用。

AI 中文摘要

设备端推理正在蓬勃发展,但这一势头几乎全部集中在语言模型上。扩散流水线内存占用大、对延迟敏感,并且需要协调嵌入器、变换器、解码器,以及通常进一步的后处理,这些后处理并不像大语言模型推理循环那样标准化。我们在性能、质量和模型占用空间之间进行权衡,以尽可能覆盖现实世界中的众多客户端设备。我们做出了三项贡献:一个嵌入翻译器,将小型文本编码器映射到大型编码器空间,以削减权重和延迟;一个可复现的扫描配方,用于在扩散流水线中导航速度/质量/内存三角;以及一个交互式设备端图像生成编辑器,在近期GPU上实现亚秒级TTFI。

英文摘要

On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a decoder, and often further postprocessing that is not as standardized as LLM inference loops are. We navigate the trade-off between performance, quality, and model footprint to reach as many client devices in the wild as possible. We make three contributions: an embedding translator that maps a small text encoder into a large encoder space to cut weight and latency; a reproducible sweep recipe for navigating the speed/quality/memory triangle in diffusion pipelines; and an interactive on-device image generation editor achieving sub-second TTFI on recent GPUs.

DOI:10.1145/3829339.3847852

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑