发表机构
Evolvent AI; The Hong Kong Polytechnic University; National University of Singapore; Columbia University(Evolvent AI; 香港理工大学; 新加坡国立大学; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对递归自我改进中搜索深度与广度的矛盾,提出Gödel Forest多智能体框架,通过共同进化树与共享记忆平衡探索,在RSIBench-Data上平均提升10.70%。
AI 中文摘要
递归自我改进(RSI)旨在通过让模型自我改进来实现复合收益。虽然现有的大多数RSI系统围绕冻结的基础模型优化外部智能体框架或提示,但数据中心的RSI通过训练智能体生成的数据直接更新模型自身的参数。然而,由于验证数据策略需要昂贵的模型训练,现有方法面临一个基本困境:单个智能体会陷入狭窄的方向,缺乏探索广度,而朴素的并行搜索或繁重的轨迹共享则牺牲了长视界的搜索深度。为了解决这一挑战,我们引入了Gödel Forest,一个多智能体框架,将递归自我改进组织为共同进化的搜索树的集成。在Gödel Forest中,每个智能体自主地生长一棵持久树,基于模型反馈深化、分支或剪枝数据策略以确保深度,而并行树探索数据空间的不同区域以扩展广度。关键的是,不是让树孤立或淹没在繁重的执行日志中,而是动态共同进化的记忆连接森林:智能体不断将其成功和失败提炼为紧凑的程序性教训,并锚定到全局排行榜。通过这种森林生态系统,一棵树中的死胡同会立即警告整个森林避开无望的路径,而经验性的突破会迅速在邻近树中播种新的探索分支。在RSIBench-Data上跨六个不同领域的评估中,Gödel Forest平均比单智能体基线高出10.70%,同时在五个任务上减少了墙钟时间。消融实验证实,共同进化的共享记忆比独立并行搜索带来+7.00%的增益,表明集体提炼是可扩展自我改进的关键。代码可在以下https URL获取。
英文摘要
Recursive self-improvement (RSI) aims to achieve compounding gains by having models improve themselves. While most existing RSI systems optimize external agent harnesses or prompts around a frozen base model, data-centric RSI directly updates the model's own parameters by training on agent-generated data. However, because validating data strategies requires expensive model training, existing methods face a fundamental dilemma: a single agent gets trapped in narrow directions and lacks exploration breadth, while naive parallel search or heavy trace sharing sacrifices long-horizon search depth. To address this challenge, we introduce G"odel Forest, a multi-agent framework that organizes recursive self-improvement as an ensemble of co-evolving search trees. In G"odel Forest, each agent autonomously grows a persistent tree, deepening, branching, or pruning data strategies based on model feedback to secure depth, while parallel trees explore distinct regions of the data space to expand breadth. Crucially, rather than leaving trees isolated or flooding them with heavy execution logs, a dynamically co-evolving memory connects the forest: agents continuously distill their successes and failures into compact procedural lessons anchored to a global leaderboard. Through this forest ecosystem, a dead-end in one tree instantly warns the whole forest against unpromising paths, while an empirical breakthrough quickly seeds new exploration branches in neighboring trees. Evaluated on RSIBench-Data across six diverse domains, G"odel Forest outperforms the single-agent baseline by an average of 10.70% while reducing wall-clock time on five tasks. Ablations confirm that co-evolving shared memory yields a +7.00% gain over independent parallel search, demonstrating that collective distillation is key to scalable self-improvement. The code is available at https://github.com/evolvent-ai/Godel-Forest.
CommentsPreprint