南贝格4.2 - 3B:以紧凑模式解锁智能体能力
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model
浏览论文内容
中文总结 AI 辅助
介绍南贝格4.2 - 3B紧凑通用智能体模型,通过特定预训练方式构建,利用多种强化学习方法提升质量与效率,在多任务和基准测试中表现出色,优于更大模型,有潜力成为紧凑本地个人助手。
中文摘要 AI 辅助
我们展示了南贝格4.2 - 3B,一个具有3B非嵌入参数的紧凑通用智能体模型。它在代码智能体、办公智能体和复杂工具使用任务中表现出色,在数学、编码和科学方面保持高度竞争力的推理能力。南贝格4.2 - 3B使用循环变压器在28T令牌上从头预训练,通过实际部署和大规模合成扩展可执行环境、任务资产和智能体支架的多样性。我们的强化学习管道应用混合模式的基于人类反馈的强化学习来提高整体模型质量并减少失败案例,长度控制推理强化学习来平衡准确性和推理效率,以及带有结果和过程奖励的智能体强化学习来稳定长期训练。广泛评估表明,南贝格4.2 - 3B在各种智能体基准测试中优于更大的模型,在推理和对齐任务中保持竞争力。与OpenClaw的性能进一步支持其作为紧凑本地个人助手的用途。
英文摘要
We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations show that Nanbeige4.2-3B outperforms larger models, including Qwen3.5-9B and Gemma4-12B, across diverse agentic benchmarks while remaining competitive on reasoning and alignment tasks. Performance with OpenClaw further supports its use as a compact local personal assistant.
发表机构
- Nanbeige LLM Lab(南北贝大语言模型实验室)
机构由 AI 辅助整理,请以论文原文为准。