发表机构
Shanghai Jiao Tong University; Carnegie Mellon University; University of Waterloo(上海交通大学; 卡内基梅隆大学; 滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RSI-Master通过实验操作系统和评审者引导的研究编排,在自主模型改进中同时避免黑客行为和策略锁定,在PostTrainBench上显著超越基线,并扩展至35B模型时超越人类开发的Instruct模型。
AI 中文摘要
递归自我改进(RSI)旨在使AI系统能够参与提升自身能力。一条具体路径是自主模型开发,其中智能体迭代探索后训练策略以改进基础模型。这一设置面临两个挑战:智能体可能通过黑客行为利用开放式实验动作,以及重复实验可能导致策略锁定,即早期方向被细化而非重新考虑。我们引入RSI-Master,它在两个层面解决这两个挑战:正则化逐步动作,避免黑客行为,并促进研究方向的良好结构化探索,避免策略锁定。RSI-Master由实验操作系统(Experiment OS)和评审者引导的研究编排(Reviewer-Guided Research Orchestration)组成,前者支持正则化的实验动作并维护持久、可追踪的实验记录,后者在动态增长的研究DAG中组织工人(Workers)和评审者(Reviewers)。工人探索多样化的研究方向,评审者比较相关实验中的证据以指导后续探索。在PostTrainBench上使用Qwen3-4B-Base,其平均得分为54.49,而最强智能体基线为46.53,黑客率为0.0%。扩展到35B模型时,RSI-Master在LiveCodeBench-v6(41.21对37.36)和SciCode上超越了人类开发的Instruct模型,并在HorizonMath上达到非零分数,该基准包含未解决的研究问题,大多数前沿模型得分接近零。
英文摘要
Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enables regularized experimental actions and maintains persistent, traceable experimental records, and Reviewer-Guided Research Orchestration, which organizes Workers and Reviewers in a dynamically growing research DAG. Workers explore diverse research directions and Reviewers compare evidence across related experiments for subsequent explorations. On PostTrainBench with Qwen3-4B-Base, it averages 54.49 versus 46.53 for the strongest agent baseline, with a 0.0\% hacking rate. Scaling to 35B model, RSI-Master surpasses the human-developed Instruct model on LiveCodeBench-v6 (41.21 vs. 37.36) and SciCode, and reaches a nonzero score on HorizonMath, a benchmark of unsolved research problems on which most frontier models score near zero.