发表机构
Scale AI; University of California, Santa Cruz; University of North Carolina at Chapel Hill(Scale AI; 加州大学圣克鲁兹分校; 北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RSI-Forge通过多智能体流水线将研究论文转化为可执行环境,生成210个跨18领域任务,验证了自我改进智能体的训练与评估可行性。
AI 中文摘要
环境是递归自我改进的基础:它们提供了智能体要解决的问题以及用于评估进展的反馈。然而,构建具有可靠评估的挑战性研究环境仍然依赖于领域专家,这限制了其规模和学科覆盖范围。我们引入了RSI-Forge,一个多智能体流水线,将已发表的论文转化为可执行的环境以用于自我改进。三个智能体协调构建、复现和审查,以产生带有自动化评估器的任务;每篇论文的方法被独立重新实现以建立基线分数。我们提供了涵盖18个领域的210个环境,其中90个由独立的领域专家进行审查。专家和智能体评审员都对改进所提供的初始解决方案的潜力以及评估器区分解决方案质量的能力给予了高度评价,而专家对捷径抵抗性、对源论文的忠实度以及单一想法是否能穷尽任务更为挑剔。为了验证其用于重复改进的有效性,我们在120个环境上对四个模型进行了连续三次尝试的评估,每次尝试继承先前的代码和笔记,而模型权重保持不变。在84%的环境中,至少有一个模型在第一次尝试后有所改进。模型还在120个环境中的68个中优于复现的论文方法,展示了超越这些基线的提升空间。转录分析显示,在95%的这些成功尝试中,工作超出了参数调整。对所得轨迹的分析表明,在这些任务上得分较低的模型探索较少,更频繁地接受小于报告标准误差的收益,并且更依赖于对开发集的调整。RSI-Forge提供了一种可扩展的方法来构建研究环境,以训练和评估自我改进的智能体。
英文摘要
Environments are the foundation of recursive self-improvement: they provide the problems agents work on and the feedback used to evaluate progress. Yet constructing challenging research environments with reliable evaluation still depends on domain experts, limiting their scale and disciplinary coverage. We introduce RSI-Forge, a multi-agent pipeline that turns published papers into executable environments for self-improvement. Three agents coordinate construction, reproduction, and review to produce tasks with automated evaluators; each paper's method is independently reimplemented to establish a baseline score. We present 210 environments across 18 fields, including 90 reviewed by independent human domain experts. Both experts and agent judges give high ratings to the potential for improving the provided starting solutions and the evaluators' ability to distinguish solution quality, whereas experts are more critical of shortcut resistance, faithfulness to the source paper, and whether a single idea can exhaust a task. To validate their use for repeated improvement, we evaluate four models over 3 successive attempts on 120 environments, with each attempt inheriting prior code and notes while model weights remain fixed. At least one model improves after the first attempt in 84% of environments. Models also outperform the reproduced paper methods in 68 of the 120 environments, demonstrating room for gains beyond these baselines. Transcript analysis identifies work beyond parameter tuning in 95% of these successful attempts. Analysis of the resulting trajectories shows that models scoring lower on these tasks explore less, more often accept gains smaller than the reported standard error, and rely more heavily on tuning to the development set. RSI-Forge provides a scalable approach to constructing research environments for training and evaluating self-improving agents.