arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33867cs.AI

R$^2$ Flow:通过递归技能进化实现递归自我改进

R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution

Mingda Zhang, Qiang Huang, Yanjin Li, Zijia Wang, Qika Lin, Xiaoying Tang, Tiesunlong Shen

首次发表
浏览论文内容

中文总结 AI 辅助

R$^2$ Flow通过递归技能进化实现递归自我改进,交替策略学习、独立验证和版本化技能库更新,在多个任务上提升准确性和编辑精度。

中文摘要 AI 辅助

基于LLM的智能体可以通过重用和修订其编排为可执行程序的技能,跨任务进行自我改进。基于流的训练适合这一循环:它按奖励比例采样程序,通过每个技能的流量为下一次库修订提供信用。有三个障碍阻碍了这种自我改进的可靠性:流训练在树状结构历史上遭受策略崩溃;非负的基于流的信用将频繁使用视为收益;库编辑依赖于策略优化的任务奖励。我们引入了R$^2$ Flow,一个递归自我改进框架,它在共享状态编排图上交替进行策略学习、独立验证和版本化技能库更新。该图合并了仅在独立步骤顺序上不同的历史,使流训练能够汇集等效执行中的证据。训练流的流共享读数,对后向策略不变,以及一个独立的符号效用排名确定要更改哪些技能,验证器证据决定编辑是否合理,残差方差平台决定何时更新。已提交的编辑重塑了下一个策略学习的图,实现了递归技能进化。在问答、数学推理、交互式决策制定和代码生成中,R$^2$ Flow在任务准确性和库编辑精度上优于启发式编排、强化学习和技能进化基线,并跨执行器转移。代码可在以下网址获取:此https URL。

英文摘要

LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obstacles stand in the way of making this self-improvement reliable: flow training suffers strategy collapse over tree-structured histories; nonnegative flow-based credit rewards frequent use as if it were benefit; and library edits rest on the task reward the policy optimizes. We introduce R$^2$ Flow, a recursive self-improvement framework that alternates policy learning, independent verification, and versioned skill-library updates on a shared-state orchestration graph. The graph merges histories that differ only in the order of independent steps, allowing flow training to pool evidence across equivalent executions. A flow-share readout of the trained flow, invariant to the backward policy, and a separate signed utility rank which skills to change, verifier evidence decides whether an edit is warranted, and a residual-variance plateau sets when to update. Committed edits reshape the graph the next policy learns on, realizing recursive skill evolution. Across question answering, mathematical reasoning, interactive decision making, and code generation, R$^2$ Flow improves task accuracy and library-edit precision over heuristic orchestration, reinforcement learning, and skill-evolution baselines, and transfers across executors. Code is available at https://github.com/beita6969/r2flow.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Fudan University(复旦大学)
  • University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • University of Oxford(牛津大学)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑