发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在通过强化学习培养大语言模型自我进化的元技能。提出MetaEvolve框架,经数据合成、进化感知强化学习等流程,在编码任务中训练模型。实验表明该框架在编码基准测试中性能出色,能有效激发通用元技能,为自主进化的人工智能发展提供路径。
AI 中文摘要
如AlphaEvolve所示,通过环境反馈进行迭代自我进化的测试时扩展显示出显著的性能提升。我们假设此类进化框架的成功取决于元技能,如环境反馈下的自我反思,这能实现有效的多轮优化,但传统训练后方法大多忽略了。为弥合这一差距,我们提出MetaEvolve框架,通过数据合成管道、进化感知强化学习和推理时进化搜索来培养这些元技能。具体而言,我们将MetaEvolve应用于编码,程序执行提供超越二进制正确性的自然连续奖励信号。基于这些信号,我们合成进化轨迹作为训练数据,训练模型并通过测试用例执行获得可验证奖励。通过在大规模代码数据上训练,旨在激发可广泛转移到缺乏此类丰富训练信号的开放式问题的通用领域无关元技能。在七个编码基准测试中,MetaEvolve在分布内任务上比最强基线绝对高出10.01%,在分布外任务上高出24.12%。在完全超出训练领域的开放式算法优化问题上,它进一步实现了46.9%的相对改进。这些结果表明,明确培养自我进化元技能为更有能力和自主自我进化的人工智能提供了一条原则性路径。
英文摘要
Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We hypothesize that the success of such evolution frameworks hinges on meta-skills, such as self-reflection with environment feedback, that enable effective multi-round refinement, yet are largely neglected by traditional post-training. To bridge this gap, we present MetaEvolve, a framework designed to develop these meta-skills via a data synthesis pipeline, evolution-aware reinforcement learning (RL), and inference-time evolutionary search. Concretely, we ground MetaEvolve in coding, where program execution provides natural, continuous reward signals beyond binary correctness. Building on these signals, we synthesize evolution trajectories as training data, each containing a current program, its fitness score (combining correctness and efficiency), and a history of prior attempts, and train the model via RL with verifiable rewards derived from test case execution. By training on large-scale code data, we aim to inspire generalizable domain-agnostic meta-skills that can transfer broadly to open-ended problems where such rich training signals are scarce. Across seven coding benchmarks, MetaEvolve outperforms the strongest baseline by 10.01% absolute on in-distribution tasks and 24.12% on out-of-distribution tasks. On open-ended algorithm optimization problems entirely outside the training domain, it further achieves a 46.9% relative improvement. These results demonstrate that explicitly cultivating self-evolution meta-skills offers a principled path toward more capable and autonomously self-evolving AI.
CommentsCOLM 2026