Ring-Zero:将零样本强化学习扩展到万亿参数以实现涌现推理
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
浏览论文内容
中文总结 AI 辅助
研究将零样本强化学习扩展到万亿参数以实现涌现推理。提出稳定高效训练管道及优化,实验发现扩展参数可提升性能、训练分阶段进行且模型产生高级认知行为,还提出评估框架,模型在推理痕迹生成上有优势。
中文摘要 AI 辅助
无需人工标注数据的可验证奖励强化学习,即零样本强化学习(zero RL),已成为引发思维链推理的强大范式。但现有研究因计算限制多局限于小模型,大规模下的训练动态和涌现能力未被探索。为应对挑战,提出稳定高效训练管道,含裁剪重要性采样等优化。实验有三个关键发现:扩展到1T参数可提升样本效率和性能上限;训练分阶段进行;模型自发产生高级认知行为。在七个数学基准测试中,Ring-2.5-1T-Zero性能有竞争力。还提出结构化评估框架,模型在生成结构化简洁推理痕迹方面有优势。希望为社区提供关于扩展行为的深入见解。
英文摘要
Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small models, leaving the training dynamics and emergent capabilities at a large scale unexplored. To meaningfully explore this frontier, we aim to elicit high-quality reasoning behaviors from the model. However, we find that naive scaling often suffers from poor readability, token redundancy, and a lack of adaptive reasoning depth. To address these challenges, we present a stable and efficient training pipeline, incorporating algorithmic and system optimizations such as clipped importance sampling, training-inference ratio correction, and mixed-precision control. Our experiments offer three key findings that validate the "bitter lesson" of scaling: (1) scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; (2) the training process progresses sequentially through an initial discovery phase followed by a sharpening phase; and (3) the model spontaneously develops advanced cognitive behaviors, including anthropomorphism, structured formatting, self-verification, parallel reasoning, and context anxiety, rendering hand-crafted heuristics redundant. Evaluated on seven mathematical benchmarks, Ring-2.5-1T-Zero achieves competitive performance. Additionally, to assess CoT quality beyond final-answer correctness, we propose a structured evaluation framework across three dimensions: comprehensibility, reproducibility, and efficiency, where our model demonstrates clear advantages in producing structured and concise reasoning traces. By sharing our observed emergent phenomena, we hope to provide the community with deeper insights into scaling behaviors, particularly at the 1-trillion scale.