arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Qwen3.8-Next架构设计:评估、效率与训练稳定性

On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu

arXiv 2608.30320首次发表:更新:

发表机构

Qwen Team(千问团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究设计了Qwen3.8-Flash-Next稀疏混合专家模型,通过混合注意力、门控残差等结构及Muon优化器,在参数、计算量大幅降低的同时,实现了更高效、更稳定且性能接近前代模型的效果。

AI 中文摘要

我们描述了Qwen3.8-Flash-Next的架构与消融实验,这是一个稀疏混合专家模型,拥有125B参数,每个token激活6B参数,另有51B参数的n-gram嵌入表存储在加速器外。在14个预训练基准上,该模型在8个基准上优于397B-A17B前代模型,其余基准上的差距最多为2.6个点,同时仅使用前代1/3的激活参数、1/3的训练token和约1/9的训练FLOPs。token混合采用门控DeltaNet(GDN)与全局注意力的分层混合结构,每4层包含1个全注意力层;在持续预训练阶段,这些全注意力层会被Qwen稀疏注意力(QSA)替代,QSA以微块粒度对上下文打分,使用压缩轻量索引器。残差流被扩展为4个分支,通过元素级门控读取,该设计称为门控残差(GR)。容量在主干外通过单个n-gram嵌入层增加,其表从主机内存预取。我们从三个维度评估每个候选变更:损失与下游基准、变更在训练、预填充和解码中的成本、对最优超参数和训练稳定性的影响。损失与下游准确率并不总是一致:扩大n-gram词汇量会持续降低损失,而下游准确率会达到饱和。该架构与Muon优化器共同将最优学习率和批量大小向上调整,使批量大小预热变得不必要,并大幅提升压力测试下的稳定性。损失、基准、效率与稳定性构成一个设计问题,联合解决后得到的方案同时更高效、更强大、更稳定。

英文摘要

We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued-pretraining time those full-attention layers are replaced by Qwen Sparse Attention (QSA), which scores context at micro-block granularity with a compressed lightweight indexer. The residual stream is widened to four branches and read through an elementwise gate, a design we call the Gated Residual (GR). Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory. We evaluate every candidate change along three axes: loss together with downstream benchmarks; the cost of the change in training, prefill and decode; and its effect on the optimal hyperparameters and training stability. Loss and downstream accuracy do not always move together: enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests. Loss, benchmarks, efficiency and stability form one design problem. Solved jointly, they yield a recipe that is simultaneously more efficient, more capable and more stable.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑