arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22577cs.AI

大规模的cMoLLM:语言模型混合的水平扩展定律

cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs

Xin Yang, Yemin Wang, Mingda Liu, Letian Li, Shuaishuai Cao, Zhengxiao He, Ryan Dong

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型训练和推理成本随模型规模线性增长的瓶颈,将MoE风格混合层 reformulate 为可变内核动态卷积,引入cMoLLM,在FineWeb训练的模型中提升了语言建模及下游任务表现,有更好流利用率等优势。

中文摘要 AI 辅助

扩展大语言模型推动了其成功,但密集型Transformer将容量和计算耦合在一起,随着模型规模增大,训练和推理成本线性增长成为瓶颈。本文旨在通过在整个大语言模型管道中采用MoE风格的混合来扩展容量。先前方法存在诸多问题,本文将MoE风格混合层重新表述为可变内核动态卷积,在此基础上引入cMoLLM。在FineWeb上训练的GPT-2风格模型中,cMoLLM在匹配计算下提高了语言建模困惑度及下游任务准确性,具有更好的流利用率、更稳定的优化和良好的扩展性。

英文摘要

Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck as models approach trillion-parameter regimes. We aim to scale capacity through MoE-style mixture throughout the LLM pipeline rather than only the FFN. Prior pipeline-level approaches include ParaScale, which introduces virtual tokens and parallel streams but incurs substantial overhead and suffers from homogenized routing and gradient collapse, and AltUp, which uses an auxiliary prediction branch but offers limited adaptivity and slow convergence. We establish that MoE-style mixture layers can be reformulated as variable-kernel dynamic convolutions, where each expert corresponds to a $1{\times}1$ convolutional kernel and routing implements input-conditioned kernel aggregation. Building on this equivalence, we introduce cMoLLM: a convolutionally gated mixture-of-LLMs that routes over end-to-end streams through fully differentiable dynamic convolution. In GPT-2-style models trained on FineWeb, cMoLLM improves language modeling perplexity and downstream GLUE and SQuAD accuracy under matched compute, with better stream utilization, more stable optimization, and favorable scaling compared to ParaScale- and AltUp-style baselines.

补充信息

↑