arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dream-RSI:通过演化世界实现递归自我改进

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Tong Zheng, Xidong Wu, Zheng Zhang, Zhankui He, Chaoyi Zhang, Benjamin Coleman, Ruoqiao Wei, Di Bai, Haolin Liu, Rui Liu, Xue Wang, Yue Zhuan, Wang-Cheng Kang, Renkai Xiang, Heng Huang, Xinwu Cheng, Yunsong Guo

arXiv 2609.14858首次发表:更新:

发表机构

University of Maryland, College Park; Google DeepMind; University of Virginia(马里兰大学帕克分校; 谷歌DeepMind; 弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Dream-RSI通过构建历史发现树的回放模拟器进行“梦境”操作,以低成本离策略反馈优化探索策略,实现递归自我改进,在多个工程领域提升发现质量并降低成本。

AI 中文摘要

递归自我改进对于自主AI智能体正变得越来越重要,其进展依赖于在复杂领域中发掘高价值解决方案。这一过程的核心驱动力是有效的探索,然而,管理和改进探索策略仍然是主要的瓶颈。当前系统面临一个根本性困境:固定策略在搜索空间扩展时无法适应,而在线策略优化则需要在长时间跨度的展开中,面对延迟且昂贵的反馈,在庞大的元搜索空间中进行导航。我们引入了Dream-RSI框架,这是一个可扩展且递归自我改进的探索框架。一个轻量级的编排层使探索变得显式且可编程,同时保持底层编码智能体不变。我们的关键洞见是,累积的发现历史可以作为已实现搜索空间上的回放模拟器。通过在由历史发现树构建的回放模拟器中进行“梦境”操作,Dream-RSI获得了即时的、低成本的离策略反馈,用于评估和改进探索策略,而无需重复调用昂贵的在线评估。改进后的策略随后被重新部署到在线环境中以推动进一步的发现,从而在自我改进的循环中不断扩展模拟器池。在算法工程、数学优化和GPU内核工程等领域,Dream-RSI实现了具有竞争力或更优的发现质量,同时在多种设置下显著降低了发现成本。

英文摘要

Recursive self-improvement is becoming essential for autonomous AI agents, whose progress depends on discovering high-value solutions across complex domains. Effective exploration drives this process, yet managing and improving exploration strategies remains a major bottleneck. Current systems face a fundamental dilemma: fixed strategies fail to adapt as search spaces scale, while online policy optimization must navigate vast meta-search spaces under delayed, expensive feedback from long-horizon rollouts. We introduce \textsc{Dream-RSI}, a framework for scalable, recursively self-improving exploration. A lightweight orchestration layer makes exploration explicit and programmable while leaving the underlying base agent unchanged. Our key insight is that accumulated discovery history can act as a replay simulator over the realized search space. By dreaming within this simulator built from historical discovery trees, \textsc{Dream-RSI} obtains immediate, low-cost off-policy feedback to evaluate and refine exploration policies without repeated, expensive online evaluation. The improved policy is then redeployed online to drive further discovery, continuously expanding the simulator pool in a self-improving loop. Across 9 tasks in 4 domains, \textsc{Dream-RSI} achieves competitive quality and improves discovery efficiency in several settings.

Comments11 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑