arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00791cs.CLcs.AI

Instella-MoE 技术报告

Instella-MoE Technical Report

Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra, Yonatan Dukler, Gowtham Ramesh, Jialian Wu, Ximeng Sun, Wen Xie, Chaojun Hou, Vikram Appia, Zhenyu Gu, Zicheng Liu, Emad Barsoum

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出完全开源的 Instella-MoE MoE 语言模型,采用创新架构与系统设计,经多阶段训练后在多项基准中表现优异,相关模型流程已开源以支撑可复现研究。

中文摘要 AI 辅助

本研究推出 Instella-MoE,这是一款完全开源的混合专家(MoE)语言模型,总参数达 160 亿,每个 token 的激活参数为 28 亿,完全从零开始在 AMD Instinct MI300X 和 MI325X GPU 上训练。Instella-MoE 结合了稀疏激活的 MoE 设计与架构及系统层面的创新,包括门控多头潜在注意力(Gated MLA)和 FarSkip-Collective 连接,实现了高效的大规模训练与推理。该模型通过多阶段流程开发,涵盖预训练、中期训练、长上下文扩展、带反馈驱动数据整理的监督微调、直接偏好优化以及多教师在线策略蒸馏强化学习。Instella-MoE 在标准预训练基准上的平均得分为 76.7,优于此前的完全开源模型,包括 OLMo-3-7B、SmolLM3-3B 和 OLMoE-1B-7B,同时在可比激活参数规模下,与 Moonlight-16B-A3B 和 Qwen3.5-4B 等开源权重 MoE 及密集基线模型表现相当。经后训练后,最终的 Think 检查点在指令跟随、推理、数学、编码和聊天基准上的平均得分为 73.2,在评估中优于具有可比或更大激活参数数量的完全开源模型及开源权重模型。为支持透明可复现的研究,我们发布了完整的 Instella-MoE 模型流程,包括模型权重、训练配置、数据混合及训练代码。这些贡献共同确立了 Instella-MoE 作为高效高性能 MoE 模型的强大完全开源基础,也为可复现研究提供了支撑。

英文摘要

In this work, we introduce Instella-MoE, a fully open Mixture-of-Experts (MoE) language model with 16 billion total parameters and 2.8 billion active parameters per token, trained entirely from scratch on AMD Instinct MI300X and MI325X GPUs. Instella-MoE combines a sparsely activated MoE design with architectural and system-level innovations, including Gated Multi-head Latent Attention (Gated MLA) and FarSkip-Collective connectivity, enabling efficient large-scale training and inference. The model is developed through a multi-stage pipeline comprising pre-training, mid-training, long-context extension, supervised fine-tuning with feedback-driven data curation, direct preference optimization, and reinforcement learning with Multi-Teacher On-Policy Distillation. Instella-MoE achieves an average score of 76.7 across standard pre-training benchmarks, outperforming prior fully open models including OLMo-3-7B, SmolLM3-3B, and OLMoE-1B-7B, while remaining competitive with open-weight MoE and dense baselines at comparable active-parameter scales, including Moonlight-16B-A3B and Qwen3.5-4B. After post-training, our final Think checkpoint achieves an average score of 73.2 across instruction-following, reasoning, math, coding, and chat benchmarks, outperforming both fully open and open-weight models with comparable or larger active parameter counts in our evaluation. To support transparent and reproducible research, we release the complete Instella-MoE model flow, including model weights, training configurations, data mixtures, and training code. Together, these contributions establish Instella-MoE a strong, fully open foundation for efficient, high-performing MoE models and reproducible research.

↑