Rufus-Air:一种开放的大语言模型后训练配方
Rufus-Air: An Open LLM Post-Training Recipe
浏览论文内容
中文总结 AI 辅助
Rufus-Air 提出一个八阶段串行后训练配方,基于开源组件和公共数据,通过难度过滤和奖励可靠性排序,显著提升 GLM-4.5-Air 性能,达到与同类开源模型竞争的水平。
中文摘要 AI 辅助
Rufus-Air 是一个在 GLM-4.5-Air-Base(106B-A12B)上开放且可复现的后训练配方,组织为八个阶段的串行流水线:SFT、推理强化学习、编码强化学习、指令遵循强化学习、通用智能体、编码智能体、搜索智能体和 RLHF。我们记录了复现该配方所需的数据、奖励设计、基础设施、阶段顺序和逐阶段结果。阶段从基础能力向高级能力推进,从硬性、可验证的奖励向更软的基于评判器的信号过渡。训练基于开源组件和公共数据,其中大部分数据按原样使用,没有新的人工标注或内部蒸馏教师。我们的主要发现是:(i)多样、高质量的 SFT 建立了强大的能力基础;(ii)难度过滤使强化学习提示保持在有效的学习范围内;(iii)奖励可靠性为排序阶段提供了实用原则;(iv)基础设施和工程选择是配方的一部分,而不仅仅是实现细节。Rufus-Air 优于官方 GLM-4.5-Air 后训练发布版本,并且与类似规模的开源模型具有竞争力。
英文摘要
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human annotation or an in-house distillation teacher. Our main findings are that (i) diverse, high-quality SFT establishes a strong capability floor; (ii) difficulty filtering keeps RL prompts within a productive learning range; (iii) reward reliability provides a practical principle for ordering stages; and (iv) infrastructure and engineering choices are part of the recipe, not just an implementation detail. Rufus-Air improves over the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models.
发表机构
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。