arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RSI-Master:结构化实验以引导自主模型改进

RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement

Yaxin Du, Xiyuan Yang, Zhifan Zhou, Yujie Ge, Cheng Wang, Jiajun Wang, Sijie Chen, Zehui Liu, Yuxin Zhang, Weicheng Gu, Julian Zhang, Zixing Lei, Siheng Chen

arXiv 2609.35561首次发表:更新:

发表机构

Shanghai Jiao Tong University; Carnegie Mellon University; University of Waterloo(上海交通大学; 卡内基梅隆大学; 滑铁卢大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RSI-Master通过实验操作系统和评审者引导的研究编排,在自主模型改进中同时避免黑客行为和策略锁定,在PostTrainBench上显著超越基线,并扩展至35B模型时超越人类开发的Instruct模型。

AI 中文摘要

递归自我改进(RSI)旨在使AI系统能够参与提升自身能力。一条具体路径是自主模型开发,其中智能体迭代探索后训练策略以改进基础模型。这一设置面临两个挑战:智能体可能通过黑客行为利用开放式实验动作,以及重复实验可能导致策略锁定,即早期方向被细化而非重新考虑。我们引入RSI-Master,它在两个层面解决这两个挑战:正则化逐步动作,避免黑客行为,并促进研究方向的良好结构化探索,避免策略锁定。RSI-Master由实验操作系统(Experiment OS)和评审者引导的研究编排(Reviewer-Guided Research Orchestration)组成,前者支持正则化的实验动作并维护持久、可追踪的实验记录,后者在动态增长的研究DAG中组织工人(Workers)和评审者(Reviewers)。工人探索多样化的研究方向,评审者比较相关实验中的证据以指导后续探索。在PostTrainBench上使用Qwen3-4B-Base,其平均得分为54.49,而最强智能体基线为46.53,黑客率为0.0%。扩展到35B模型时,RSI-Master在LiveCodeBench-v6(41.21对37.36)和SciCode上超越了人类开发的Instruct模型,并在HorizonMath上达到非零分数,该基准包含未解决的研究问题,大多数前沿模型得分接近零。

英文摘要

Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enables regularized experimental actions and maintains persistent, traceable experimental records, and Reviewer-Guided Research Orchestration, which organizes Workers and Reviewers in a dynamically growing research DAG. Workers explore diverse research directions and Reviewers compare evidence across related experiments for subsequent explorations. On PostTrainBench with Qwen3-4B-Base, it averages 54.49 versus 46.53 for the strongest agent baseline, with a 0.0\% hacking rate. Scaling to 35B model, RSI-Master surpasses the human-developed Instruct model on LiveCodeBench-v6 (41.21 vs. 37.36) and SciCode, and reaches a nonzero score on HorizonMath, a benchmark of unsolved research problems on which most frontier models score near zero.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑