arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12176cs.AI

通过多智能体自监督实现递归自我改进

Recursive Self-Improvement through Multi-Agent Self-Supervision

Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Somayeh Sojoudi, Matei Zaharia, Yujin Tang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对模型递归自我改进的监督瓶颈,提出多智能体自监督(MASS)方法,经Qwen3.6-27B验证,可提升开放式基准性能并提高训练效率。

中文摘要 AI 辅助

模型在无法验证的任务(如开放式研究)上的递归自我改进(RSI)面临监督瓶颈:当模型输出超出人类专家可可靠评估的范围时,模型本身(优化对象)成为最佳可用优化器和评估器。然而,单一模型实例在这种同构循环中难以批判和改进自身的复杂推理。为解决该问题,我们提出多智能体自监督(MASS),一种交替进行进化工作流优化和对自生成轨迹进行监督微调的RSI方法。基于早期发现(多智能体拓扑在复杂推理中表现出色),MASS提示单一基础模型迭代提出、执行和自我评估多智能体工作流。在结构护栏约束下的进化搜索中,模型优化这些类计算图的编排,为给定任务发现最有效的不同角色和信息路由。在使用Qwen3.6-27B的两个MASS循环中,模型在四个开放式公共基准上每输出token的性能提升1.2-1.6倍。由于改进后的模型随后成为更好的优化器和评估器,这种交替框架实现了模型能力的持续递归自举。此外,多智能体轨迹的训练效率更高:在其上训练的学生模型,性能优于在多1.4倍训练token上训练的单智能体学生模型。这些发现表明,从多智能体轨迹中联合学习编排和有界子智能体执行,可为RSI提供有效信号。

英文摘要

Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance struggles to critique and improve its own complex reasoning under this homogeneous loop. To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories. Guided by early findings that multi-agent topologies excel at complex reasoning, MASS prompts a single base model to iteratively propose, execute, and self-evaluate multi-agent workflows. Through an evolutionary search constrained by structural guardrails, the model optimizes these computational-graph-like orchestrations, discovering the most effective distinct roles and information routing for a given task. Over two MASS cycles with Qwen3.6-27B, the model achieves 1.2-1.6x higher performance per output tokens on four open-ended public benchmarks. Because the improved model subsequently acts as a better optimizer and evaluator, this alternating framework enables a continuous, recursive bootstrapping of the model's capabilities. Moreover, multi-agent traces are also more training-efficient: a student trained on them outperforms a single-agent student trained on 1.4x more training tokens. These findings suggest that jointly learning orchestration and bounded subagent execution from multi-agent trajectories can provide an effective signal for RSI.

发表机构

  • UC Berkeley(加州大学伯克利分校)
  • Sakana AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑