AI 中文总结
ROSS通过选择性监督从历史自生成轨迹中重新学习,在多个任务上提升性能,无需额外策略轨迹。
AI 中文摘要
大型语言模型的后训练通过强化学习和在线策略蒸馏生成自生成轨迹,然而,一旦策略更新,这些经验往往被视为过时。历史轨迹可以与后续策略保持兼容,同时保留策略不再可靠表达的行为。然而,这些轨迹也可能包含错误、放弃的尝试以及不应模仿的冗余动作,这促使需要更细粒度的选择性监督。我们提出了ROSS(通过选择性监督从自生成轨迹中重新学习),该方法保留完整的历史轨迹作为上下文,而仅对选定的模型生成续段应用损失。在领域特定的强化学习、多教师在线策略蒸馏以及智能体强化学习中,ROSS一致地改进了上游检查点,并在数学、代码生成、指令遵循和软件工程方面优于基线。在Qwen3.6-35B-A3B上,ROSS将六基准MOPD平均值从58.40%提升至62.20%,SWE-bench Verified从64.20%提升至68.40%。这些结果表明,自轨迹训练留下了可重复使用的行为经验,通过离线监督微调(SFT)可以进一步获得收益,而无需额外的策略轨迹。
英文摘要
Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.
Comments21 pages, 6 figures, 10 tables