发表机构
University of the Chinese Academy of Sciences; Baidu; Tsinghua University; Institute of Automation, Chinese Academy of Sciences; Shanghai University(中国科学院大学; 百度; 清华大学; 中国科学院自动化研究所; 上海大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AGAR将程序进化形式化为马尔可夫决策过程,以模块化前缀为动作,使强化学习组件可独立替换,在19个任务上优于强基线,增益集中于竞争性编程。
AI 中文摘要
给定一个任务和一个评估器,语言模型可以重写一个候选程序,而搜索循环决定哪些重写能够存活,这为算法发现提供了一条实用途径。但该循环由五个手工设定的常量控制:选择哪个父代、变异强度如何、如何保持多样性、记住什么,以及一个标量分数,该分数从不说明程序的哪一部分获得了该分数。强化学习已经为每一个常量提供了相应的估计器。障碍在于,程序进化通常不会被表述为一个决策过程。我们将其形式化为一个马尔可夫决策过程,其动作是模型所依赖的模块化前缀,而非模型生成的程序。这样,信用分配、价值估计、自适应探索和经验记忆就可以分别附加到不同的组件上。AGAR(算法生成即强化学习)提供了由此产生的基板:任何估计器都可以在不改变控制器的情况下被替换或关闭,使得迁移可以一次一个机制地进行审计,且无需对后端模型进行梯度训练。在统一测试框架下,跨越19个任务、两个后端和三个随机种子,AGAR在大多数任务上优于两个已发表基线中较强的那个,其增益集中在竞争性编程任务族中。该形式化还为先前的相关工作提供了一种可检验的解读:这些系统隐式地采用零折扣,并非出于选择,而是因为适应度是相对于个体而言的外生因素,而非相对于后继者的回报,这使得折扣因子无所作用。
英文摘要
Given a task and an evaluator, a language model can rewrite a candidate program while a search loop decides which rewrites survive, offering a practical route to algorithm discovery. But that loop is governed by five constants set by hand: which parent to select, how hard to mutate, how to keep diversity, what to remember, and a scalar score that never says which part of the program earned it. Reinforcement learning already has an estimator for each. The obstacle is that program evolution is not usually written down as a decision process. We formalize it as a Markov decision process whose action is the modular prefix the model is conditioned on, rather than the program it emits. Credit assignment, value estimation, adaptive exploration, and experience memory can then attach to distinct components. AGAR (Algorithm Generation As RL) provides the resulting substrate: any estimator can be replaced or switched off without changing the controller, making the transfer auditable one mechanism at a time, with no gradient training of the backend model. Across 19 tasks, two backends, and three seeds under one harness, AGAR improves on the stronger of two published baselines on most tasks, with gains concentrated in the competitive-programming family. The formalization also yields a checkable reading of prior work: these systems are implicitly zero-discount, not by choice, but because fitness is exogenous to an individual rather than a return over successors, leaving a discount factor nothing to act on.