发表机构
Qwen Team(千问团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究设计了Qwen3.8-Flash-Next稀疏混合专家模型,通过混合注意力、门控残差等结构及Muon优化器,在参数、计算量大幅降低的同时,实现了更高效、更稳定且性能接近前代模型的效果。
AI 中文摘要
我们描述了Qwen3.8-Flash-Next的架构与消融实验,这是一个稀疏混合专家模型,拥有125B参数,每个token激活6B参数,另有51B参数的n-gram嵌入表存储在加速器外。在14个预训练基准上,该模型在8个基准上优于397B-A17B前代模型,其余基准上的差距最多为2.6个点,同时仅使用前代1/3的激活参数、1/3的训练token和约1/9的训练FLOPs。token混合采用门控DeltaNet(GDN)与全局注意力的分层混合结构,每4层包含1个全注意力层;在持续预训练阶段,这些全注意力层会被Qwen稀疏注意力(QSA)替代,QSA以微块粒度对上下文打分,使用压缩轻量索引器。残差流被扩展为4个分支,通过元素级门控读取,该设计称为门控残差(GR)。容量在主干外通过单个n-gram嵌入层增加,其表从主机内存预取。我们从三个维度评估每个候选变更:损失与下游基准、变更在训练、预填充和解码中的成本、对最优超参数和训练稳定性的影响。损失与下游准确率并不总是一致:扩大n-gram词汇量会持续降低损失,而下游准确率会达到饱和。该架构与Muon优化器共同将最优学习率和批量大小向上调整,使批量大小预热变得不必要,并大幅提升压力测试下的稳定性。损失、基准、效率与稳定性构成一个设计问题,联合解决后得到的方案同时更高效、更强大、更稳定。
英文摘要
We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued-pretraining time those full-attention layers are replaced by Qwen Sparse Attention (QSA), which scores context at micro-block granularity with a compressed lightweight indexer. The residual stream is widened to four branches and read through an elementwise gate, a design we call the Gated Residual (GR). Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory. We evaluate every candidate change along three axes: loss together with downstream benchmarks; the cost of the change in training, prefill and decode; and its effect on the optimal hyperparameters and training stability. Loss and downstream accuracy do not always move together: enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests. Loss, benchmarks, efficiency and stability form one design problem. Solved jointly, they yield a recipe that is simultaneously more efficient, more capable and more stable.