发表机构
CINECA(意大利国家计算应用研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍了基于9万亿令牌多语言数据预训练的100亿参数推理语言模型Domyn-Small,经持续预训练、监督微调、强化学习等训练后管道,实现双模式推理,在与对等模型对比中平衡了准确性与效率,还发布了相关权重及开源框架。
AI 中文摘要
我们介绍了Domyn-Small,这是一个在麻省理工学院许可下发布的拥有100亿参数的开放权重推理语言模型。Domyn-Small是在9万亿个令牌的多语言数据上进行初始预训练阶段的产物,随后是用于推理、指令跟随和上下文扩展的训练后管道。对于后者,我们进行了持续预训练(CPT)阶段,将原生上下文窗口翻倍至32K令牌,然后是带有数学聚焦退火运行的监督微调(SFT)。最后,强化学习(RL)阶段包括具有可验证奖励的广义信赖域策略优化(GRPO)、近端策略优化(DPO)以及跨越五个任务领域的多环境GRPO阶段。32K令牌的原生上下文在推理时通过YaRN扩展到128K,并且聊天模板切换实现了双模式推理。与70亿至100亿类别的对等模型相比,Domyn-Small实现了强大的准确性-效率平衡。我们还发布了权重和训练后方法,以及Domyn Swarm(Apache~2.0),这是一个在此项目中开发并在整个工作中使用的用于在HPC集群上进行可扩展大语言模型推理的开源框架。
英文摘要
We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under the MIT license. Domyn-Small is the product of an initial pre-training phase on 9 trillion tokens multilingual data, followed by a post-training pipeline for reasoning, instruction following, and context extension. For the latter, we performed a Continued Pre-Training (CPT) phase that doubles the native context window to 32K tokens, followed by SFT with a math-focused annealing run. Finally, the RL phase includes GRPO with verifiable rewards, DPO, and a multi-environment GRPO stage spanning five task domains: mathematics, code, multiple-choice QA, instruction-following, and tool calling. The 32K-token native context extends to 128K at inference via YaRN, and a chat-template toggle enables dual-mode reasoning. Against peer models in the 7--10B class (Qwen3.5-9B, OLMo-3-7B-Think, Nemotron-Nano-8B, Ministral-3-8B), Domyn-Small achieves a strong accuracy-efficiency balance: it produces roughly one-third as many tokens as Qwen3.5-9B and approximately 35% of OLMo-3-7B-Think's token budget on core reasoning benchmarks, while delivering strong instruction-following (IFEval 79.9) and competitive science reasoning (GPQA-Diamond 50.0). We release the weights and the post-training recipe alongside Domyn Swarm (Apache~2.0), an open-source framework for scalable LLM inference on HPC clusters developed during this program and used throughout this work.
Comments27 pages, 5 figure