arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RSIGym:用于递归自我改进的灵活环境

RSIGym: A Flexible Environment for Recursive Self-Improvement

Fanqing Meng, Lingxiao Du, Haocheng Lu, Qiguang Chen, Ziqi Zhao, Zijian Wu, Jiayuan Zhuo, Mengkang Hu, Michael Qizhe Shieh

arXiv 2610.10310首次发表:更新:

发表机构

Evolvent AI; National University of Singapore(Evolvent AI; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RSIGym基于一切皆服务理念,为递归自我改进提供统一环境,通过可复用服务和联合优化,使Opus 5在五个基准上取得最高RSI-Index 0.4809,显著提升性能。

AI 中文摘要

递归自我改进要求将已接受的变更带入后续改进周期,而研究智能体提出的变更也需要大量的研究基础设施。现有设置往往让智能体重建常规基础设施,或将探索限制在单个组件上。我们引入了RSIGym,一个基于“一切皆服务”(EaaS)的智能体原生研究环境。RSIGym通过可复用服务提供训练、推理、回滚、评估和沙盒执行,并通过共享预算和权限控制支持数据、框架和联合改进轨道。这种设计使智能体能够在同一环境中调查单个干预措施,并联合优化数据、训练设置和执行框架。我们将RSI-Index定义为在覆盖软件工程、终端交互、数学、科学推理和基于技能的任务的五个基准上,剩余性能差距被闭合的平均比例。在独立的联合运行中比较六个前沿研究模型,Opus 5在每次基准运行的500美元平台服务预算下实现了最高的RSI-Index 0.4809。其选择的系统改进了所有五个基准,将SWE-bench Verified从17.67%提升至50.33%,将AIME从31.67%提升至97.78%。额外实验考察了DSH框架细化、预算变化和受限网络访问,而记录的轨迹揭示了智能体如何诊断失败和选择候选方案。我们开源了完整的RSIGym代码库和结果,以支持可复现性和进一步研究。

英文摘要

Recursive self-improvement requires carrying accepted changes into later improvement cycles, while studying agent-proposed changes also requires substantial research infrastructure. Existing settings often leave agents to rebuild routine infrastructure or restrict exploration to individual components. We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS). RSIGym exposes training, inference, rollout, evaluation, and sandbox execution through reusable services, with shared budget and permission controls supporting Data, Harness, and Joint improvement tracks. This design enables agents to investigate individual interventions and jointly optimize data, training settings, and execution harnesses within the same environment. We define RSI-Index as the mean fraction of the remaining performance gap closed across five benchmarks covering software engineering, terminal interaction, mathematics, scientific reasoning, and skill-based tasks. Comparing six frontier research models in independent Joint runs, Opus 5 achieves the highest RSI-Index of 0.4809 under a $500 platform-service budget per benchmark run. Its selected systems improve all five benchmarks, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. Additional experiments examine DSH-harness refinement, budget variation, and restricted network access, while recorded trajectories reveal how agents diagnose failures and select candidates. We open-source the full RSIGym codebase and results to support reproducibility and further research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑