arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Motif 3:技术报告

Motif 3: Technical Report

Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung, Gihun Cho, Minsu Ha, Sangho Kang, Beomgyu Kim, Dongseok Kim, Jangwoong Kim, Taehyun Kim, Taewhan Kim, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Dongpin Oh, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Hanbin Jung, Changjin Kang, Minjae Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee, Yeongjae Park, Bokki Ryu

arXiv 2608.09119首次发表:更新:

发表机构

Motif Technologies(Motif科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出仅解码器的MoE语言模型Motif 3,采用GDLA等技术,经万亿token预训练及多阶段后训练,在长上下文、推理等任务上性能与领先开源模型相当。

AI 中文摘要

我们推出了Motif 3,这是一种仅解码器的混合专家(MoE)语言模型,总参数达3140亿,每个token激活132亿参数。每个稀疏MoE层包含384个路由专家,每个token选择8个,这种细粒度稀疏性在限制计算量的同时提供了充足的专家容量。Motif 3基于分组差分潜在注意力(GDLA)构建,该技术将分组差分注意力与多头潜在注意力的压缩键值表示相融合。该架构还采用了改进的流形约束超连接、专家特定的PolyNorm激活以及多token预测,以提升优化稳定性、专家专业化程度和推理效率。我们在约12.5万亿个token上对Motif 3进行预训练,这些token涵盖网页文档、STEM、代码、数学、多语言内容及领域专用语料库。专家平衡与数值稳定技术支持大规模稳定训练,而选择性MXFP8计算与通信、内存高效融合内核以及窗口感知上下文并行性则支持上下文长度达256K token的训练。我们的后训练流程结合了通用监督微调、6个通过强化学习训练的专家教师、1个通过监督微调训练的软件工程教师,以及多教师在线策略蒸馏。最终得到的统一模型整合了推理、编码、工具使用、专业工作、长上下文理解、校准弃权(不执行)和指令遵循等互补能力。在广泛的评估套件中,Motif 3展现出与领先开源权重模型相当的性能,在长 horizon 智能体任务、数学推理、科学知识及幻觉敏感评估中表现优异。

英文摘要

We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction to improve optimization stability, expert specialization, and inference efficiency. We pretrain Motif 3 on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora. Expert-balancing and numerical-stabilization techniques support stable training at scale, while selective MXFP8 computation and communication, memory-efficient fused kernels, and window-aware context parallelism enable training with context lengths up to 256K tokens. Our post-training pipeline combines general supervised fine-tuning, six specialist teachers trained with reinforcement learning, a software-engineering teacher trained with supervised fine-tuning, and Multi-teacher On-Policy Distillation. The resulting unified model consolidates complementary capabilities in reasoning, coding, tool use, professional work, long-context understanding, calibrated abstention, and instruction following. Across a broad evaluation suite, Motif 3 demonstrates competitive performance against leading open weight models, including strong results on long-horizon agentic tasks, mathematical reasoning, scientific knowledge, and hallucination-sensitive evaluation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑