arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16184cs.LG

PagedWeight:通过动态质量感知权重量化实现高效的混合专家语言模型服务

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization

Yuchen Yang, Yifan Zhao, Anisha Dasgupta, Sasa Misailovic

首次发表
浏览论文内容

中文总结 AI 辅助

研究KV缓存密集场景下MoE模型权重与缓存的矛盾,提出PagedWeight方法,通过动态量化权重平衡精度与缓存大小,在多个内存敏感场景中改进质量-内存权衡,实现高精度、高内存节省及吞吐量提升。

中文摘要 AI 辅助

混合专家(MoE)是一类流行的大语言模型,具有高效率和准确性。然而,在KV缓存密集的服务场景中,MoE模型权重的GPU内存需求与不断增长的KV缓存之间存在矛盾。我们提出了PagedWeight,一种用于MoE语言模型服务的新型管理方法,它在运行时动态量化MoE模型的权重,并在专家权重精度与KV缓存大小之间取得平衡。PagedWeight揭示并有效应对了模型任务精度、内存消耗和吞吐量/延迟之间复杂的权衡。在多个内存敏感的MoE服务场景中,PagedWeight相对于现有的量化基线改进了质量-内存权衡。PagedWeight实现了与FP16相当的精度,GPU内存节省高达72.0%,吞吐量提高1.94倍,并且在类似的内存预算下,相对于量化方法,质量提高了39.3%,吞吐量损失最多为4.1%。

英文摘要

Mixture-of-Experts (MoE) is a popular class of large language models (LLMs), offering high efficiency and accuracy. However, in KV-cache-intensive serving scenarios, MoEs often exhibit a tension between the GPU memory requirements of the model weights and the growing KV cache. We propose PagedWeight, a novel management method for MoE LLM serving that dynamically quantizes MoE model's weights at runtime and balances expert-weight precision with the KV cache sizes. PagedWeight exposes and effectively navigates the complex tradeoff between the model's task accuracy, memory consumption, and throughput/latency. Across several memory-sensitive MoE serving scenarios, PagedWeight improves the quality-memory tradeoff over several existing quantization baselines. PagedWeight achieves FP16-equivalent accuracy with up to 72.0% GPU memory savings and 1.94$\times$ throughput improvement, and improves quality over quantization methods by up to 39.3% at a similar memory budget with at most 4.1% throughput loss.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑