arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21693cs.LG

对混合专家大型语言模型中的可组合压缩技术进行基准测试

Benchmarking Composable Compression Techniques in Mixture-of-Experts LLMs

Afsara Benazir, Chen Chen, Rongxiao Qu, Jiabo Huang, Jingtao Li, Lingjuan Lyu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出MoEXBench基准,系统评估混合专家LLM的可组合压缩技术,发现压缩方法间存在非平凡相互作用,为MoE模型部署提供实用的准确率-内存-延迟比较方案。

中文摘要 AI 辅助

混合专家(MoE)大型语言模型通过稀疏激活高效扩展模型容量,但其庞大的专家参数占用、路由不平衡以及长上下文键值缓存(KV-cache)的增长,使得在商用硬件上部署十分困难。实际部署通常需要堆叠多种压缩技术:专家剪枝去除冗余专家,权重量化降低模型内存占用,KV缓存压缩减少长上下文内存压力。然而,这些技术通常是单独评估的,尚不清楚它们在实际部署流水线中共同应用时如何相互作用。在本研究中,我们提出MoEXBench,这是一个用于将可组合MoE压缩作为端到端部署工作流进行评估的系统基准。MoEXBench研究了10个总参数规模从300亿到2350亿的MoE模型,涵盖标准注意力、混合线性注意力和滑动窗口注意力架构。它评估了20%-50%的专家剪枝率、1到16位的权重量化方案,以及多种KV缓存精度设置,这些设置既单独应用也组合应用。MoEXBench引入了一个八模块评估套件,共同测量可组合压缩质量、工作负载和架构鲁棒性、剪枝/量化/KV缓存敏感性,以及在商用硬件上的部署效率。我们的结果揭示了压缩方法之间存在重大的相互作用:可组合压缩无法从单独的技术中预测,仅压缩率无法可靠预测质量损失或运行时增益,专家剪枝是主要的性能下降来源,而平均质量会掩盖特定工作负载和架构的故障。通过发布标准化模块分数、压缩产物和可复现脚本,MoEXBench支持在MoE家族和硬件后端之间进行实用的准确率-内存-延迟比较。

英文摘要

Mixture-of-Experts (MoE) LLMs scale model capacity efficiently through sparse activation, but their large expert parameter footprint, routing imbalance, and long-context KV-cache growth make deployment difficult on commodity hardware. Practical deployment often requires stacking multiple compression techniques: expert pruning removes redundant experts, weight quantization lowers model memory footprint, and KV-cache compression reduces long-context memory pressure. However, these techniques are typically evaluated in isolation, leaving open how they interact when applied together in realistic deployment pipelines. In this work, we present MoEXBench, a systematic benchmark for evaluating composable MoE compression as an end-to-end deployment workflow. MoEXBench studies 10 MoE models ranging from 30B to 235B total parameters across standard-attention, hybrid linear-attention, and sliding window attention architectures. It evaluates 20%-50% expert pruning rates, 1 to 16 bit weight-quantization schemes, and multiple KV-cache precision settings, applied both individually and in combination. MoEXBench introduces an eight-module evaluation suite that jointly measures composable-compression quality, workload and architecture robustness, pruning/quantization/KV cache sensitivity, and deployment efficiency on commodity hardware. Our results reveal non-trivial interactions among compression methods: composable compression cannot be predicted from standalone techniques, compression rate alone does not reliably predict quality loss or runtime gain, expert pruning is the dominant degradation source, and average quality can hide workload and architecture-specific failures. By releasing normalized module scores, compressed artifacts, and reproducible scripts, MoEXBench enables practical accuracy-memory-latency comparison across MoE families and hardware backends.

发表机构

  • University of Virginia(弗吉尼亚大学)
  • Sony AI(索尼人工智能)

机构由 AI 辅助整理,请以论文原文为准。

↑