发表机构
Motif Technologies(Motif科技公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出仅解码器的MoE语言模型Motif 3,采用GDLA等技术,经万亿token预训练及多阶段后训练,在长上下文、推理等任务上性能与领先开源模型相当。
AI 中文摘要
我们推出了Motif 3,这是一种仅解码器的混合专家(MoE)语言模型,总参数达3140亿,每个token激活132亿参数。每个稀疏MoE层包含384个路由专家,每个token选择8个,这种细粒度稀疏性在限制计算量的同时提供了充足的专家容量。Motif 3基于分组差分潜在注意力(GDLA)构建,该技术将分组差分注意力与多头潜在注意力的压缩键值表示相融合。该架构还采用了改进的流形约束超连接、专家特定的PolyNorm激活以及多token预测,以提升优化稳定性、专家专业化程度和推理效率。我们在约12.5万亿个token上对Motif 3进行预训练,这些token涵盖网页文档、STEM、代码、数学、多语言内容及领域专用语料库。专家平衡与数值稳定技术支持大规模稳定训练,而选择性MXFP8计算与通信、内存高效融合内核以及窗口感知上下文并行性则支持上下文长度达256K token的训练。我们的后训练流程结合了通用监督微调、6个通过强化学习训练的专家教师、1个通过监督微调训练的软件工程教师,以及多教师在线策略蒸馏。最终得到的统一模型整合了推理、编码、工具使用、专业工作、长上下文理解、校准弃权(不执行)和指令遵循等互补能力。在广泛的评估套件中,Motif 3展现出与领先开源权重模型相当的性能,在长 horizon 智能体任务、数学推理、科学知识及幻觉敏感评估中表现优异。
英文摘要
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction to improve optimization stability, expert specialization, and inference efficiency. We pretrain Motif 3 on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora. Expert-balancing and numerical-stabilization techniques support stable training at scale, while selective MXFP8 computation and communication, memory-efficient fused kernels, and window-aware context parallelism enable training with context lengths up to 256K tokens. Our post-training pipeline combines general supervised fine-tuning, six specialist teachers trained with reinforcement learning, a software-engineering teacher trained with supervised fine-tuning, and Multi-teacher On-Policy Distillation. The resulting unified model consolidates complementary capabilities in reasoning, coding, tool use, professional work, long-context understanding, calibrated abstention, and instruction following. Across a broad evaluation suite, Motif 3 demonstrates competitive performance against leading open weight models, including strong results on long-horizon agentic tasks, mathematical reasoning, scientific knowledge, and hallucination-sensitive evaluation.