arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过L0正则化混合专家加速稠密大语言模型

Accelerating Dense LLMs via L0-regularized Mixture-of-Experts

Zhenyu Zhang, Jiudong Yang, Zhaowen Tao, Meng Chen

arXiv 2609.21672首次发表:更新:

发表机构

YZW; FuTu AI; Wise AI(YZW; 富途人工智能; Wise AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出L0-MoE,利用L0正则化的轻量级混合专家方法,在不明显损失性能的前提下加速稠密大语言模型,实现高达2.5倍推理加速,超越现有加速基线。

AI 中文摘要

大语言模型(LLMs)性能强劲,但推理速度慢且成本高昂。现有加速方法往往导致明显的性能下降,而混合专家(MoE)模型则需要大量的计算资源。在本文中,我们提出L0-MoE,一种使用L0正则化的轻量级MoE方法,可在几乎不损失性能的情况下加速稠密LLMs。我们的方法引入了一个簇混淆矩阵用于领域感知的数据集整理,并应用动态批处理以实现高效训练。实验表明,与稠密模型相比,L0-MoE实现了高达2.5倍的加速,同时保持了有竞争力的性能,超越了现有的LLM加速基线。

英文摘要

Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources. In this paper, we propose L0-MoE, a lightweight MoE approach using L0-regularization to accelerate dense LLMs nearly without performance loss. Our method introduces a cluster confusion matrix for domain-aware dataset curation and applies dynamic batching for efficient training. Experiments show that L0-MoE achieves up to 2.5x speedup over dense models while maintaining competitive performance, outperforming existing LLM acceleration baselines.

Journal refProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑