发表机构
University of the Chinese Academy of Sciences; Baidu; Tsinghua University; Institute of Automation, Chinese Academy of Sciences; Peking University; Shanghai University; University of Bristol(中国科学院大学; 百度; 清华大学; 中国科学院自动化研究所; 北京大学; 上海大学; 布里斯托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出RLDiscover框架,通过LLM驱动的渐进式协同进化与概率评估,自动改进深度强化学习算法,在多个基准上显著提升回报并发现可解释的算法组件。
AI 中文摘要
LLM引导的程序进化已在数学和计算优化领域取得发现,这提升了强化学习(RL)算法自我进化以改进智能体学习方式的前景。然而,实现这一前景面临两个障碍。对耦合算法组件的联合搜索难以扩展:同时更改会破坏学习过程,而孤立更改则忽略了它们之间的依赖关系。评估候选算法还需要昂贵的训练,且适应度在不同随机种子间仍存在不确定性。我们引入了RLDiscover,一个用于无模型深度强化学习算法自我进化的框架。渐进式协同进化从针对性的组件编辑推进到联合进化,而渐进式概率评估通过分阶段训练和重复评估来平衡搜索广度与评估保真度。在SAC、PPO和DQN上跨四个基准套件的实验表明,平均回报显著提升,每族中位增益为32%-84%,峰值回报比率约为近零基线的363倍。这些增益包括从失败学习到成功任务完成的转变,且当进化从更强的开源实现开始时,改进仍然持续存在。在测得的SAC运动任务运行中,评估所用计算量约为完全评估同一候选池所需估算计算量的十五分之一。值得注意的是,独立搜索反复发现可解释的自适应鲁棒损失、进度相关价值目标和运行统计的组合,且选定的程序可迁移到未见任务。这些发现指向自我进化在AI中的更广泛作用:发现可解释的算法以改进智能体的学习方式。
英文摘要
LLM-guided program evolution has enabled discoveries in mathematics and computational optimization, raising the prospect of reinforcement learning (RL) algorithms that self-evolve to improve how agents learn. However, realizing this prospect faces two obstacles. Joint search over coupled algorithmic components is difficult to scale: simultaneous changes can disrupt learning, while isolated changes overlook their dependencies. Evaluating candidate algorithms also requires costly training, with fitness remaining uncertain across random seeds. We introduce RLDiscover, a framework for the self-evolution of model-free deep RL algorithms. Progressive Co-Evolution advances from targeted component edits to joint evolution, while Progressive Probabilistic Evaluation balances search breadth and evaluation fidelity through staged training and repeated evaluation. Experiments across SAC, PPO, and DQN on four benchmark suites show substantial improvements in mean return, with per-family median gains of 32%-84% and a peak return ratio of approximately 363x over a near-zero baseline. These gains include transitions from failed learning to successful task completion, and improvements persist when evolution starts from stronger open-source implementations. On measured SAC locomotion runs, evaluation uses approximately one-fifteenth the estimated compute required to fully evaluate the same candidate pool. Remarkably, independent searches repeatedly discover interpretable combinations of adaptive robust losses, progress-dependent value targets, and running statistics, with selected programs transferring to unseen tasks. These findings point toward a broader role for self-evolution in AI: discovering interpretable algorithms that improve how agents learn.