发表机构
Zhongguancun Academy; Zhongguancun Institute of Artificial Intelligence(中关村学院; 中关村人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ZGCM-1是一个完全开放的7B参数基础模型,通过内部思考与外部工具结合及高效训练方案,在数学推理和智能体搜索上媲美超大模型,并开源全流程资源。
AI 中文摘要
在这项工作中,我们提出了ZGCM-1,一个完全开放的70亿参数稠密基础模型,从零开始训练,在数据、系统和算法效率上实现了极致优化。ZGCM-1基于一个核心前提:紧凑模型无法被动记忆开放网络,但可以通过将深思熟虑的内部思考与主动的外部工具使用相结合,来克服参数容量限制。为了在256K上下文长度下支持这一范式,我们开发了一套端到端的高效开放训练方案:架构与系统协同设计:交错的门控滑动窗口注意力与全注意力,以及稳定的FP8 Muon优化器;渐进式课程与MDP中期训练:在16K、64K和256K上进行上下文扩展,并将交互轨迹重构为马尔可夫决策过程。此外,我们建立了一个AI原生的研发工作流,其中智能体群自主管理集群操作、数据整理和快速诊断评估。大量评估表明,ZGCM-1-7B在通用基准测试上与7B模型家族具有竞争力。在几个具有挑战性的数学推理和智能体搜索套件上,它与规模大数个数量级的先进模型(如Qwen3-235B-A22B和GLM-5.1)保持竞争力。我们还表明,我们的预训练设计在16K预训练时间到损失方面提供了约4.2倍的效率提升。在整个开发生命周期中,我们提炼了八项可操作的实证发现,涵盖架构扩展、SFT质量剪枝、长上下文泛化和智能体协同训练动态。为促进社区研究,我们开源了预训练、中期训练和后训练阶段的模型权重、中间检查点、训练代码、各阶段数据和数据配方,以及W&B日志。
英文摘要
In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use. To support this paradigm across a 256K context, we develop an end-to-end, high-efficiency open training recipe: Architecture & System Co-design: interleaved gated sliding-window and full attention, and a stable FP8 Muon optimizer; Progressive Curriculum & MDP Mid-Training: context scaling across 16K, 64K, and 256K, and the reformulation of interaction traces into Markov Decision Processes. Furthermore, we establish an AI-native R&D workflow where agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation. Extensive evaluations show that ZGCM-1-7B is competitive across 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1. We also show that our pre-training design offers a ~4.2x efficiency improvement in 16K pre-training time-to-loss. Across the full development lifecycle, we distill eight actionable empirical findings-spanning architectural scaling, SFT quality pruning, long-context generalization, and agentic co-training dynamics. To facilitate community research, we open-source model weights from the pre-training, mid-training, and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs.