Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
Nemotron-Cascade:基于级联强化学习的通用推理模型规模化
机构 * NVIDIA(英伟达)
AI总结 本文提出级联域强化学习(Cascade RL)方法,开发出Nemotron-Cascade模型,可在指令和深度思考模式间切换,性能不逊于纯思考模型,同时提升推理能力并保持跨领域基准性能。
Comments We publicly release the Nemotron-Cascade models and the full collection of training data at: https://huggingface.co/collections/nvidia/nemotron-cascade