发表机构
IT University of Copenhagen; University of Pisa; Sakana AI(哥本哈根信息技术大学; 比萨大学; 萨卡纳人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将神经进化(NE)作为持续强化学习的替代方案,发现进化策略(ES)在稳定性-可塑性权衡上优于强化学习变体,且权重空间邻域重叠与此权衡相关,确立了NE的竞争力。
AI 中文摘要
尽管已有许多研究探讨了持续任务变化下强化学习(RL)中可塑性损失的原因与补救措施,但尚无任何RL方法能持续地在适应与遗忘之间取得良好平衡。在此,我们转向一种替代性优化范式——神经进化(NE):这类算法通过在权重空间中直接搜索,对神经网络群体进行变异与选择。在广泛的环境与环境变化中,使用从数百个参数到百万参数网络的策略,我们将进化策略(ES)和遗传算法(GA)与最先进的持续RL变体及基于群体的RL进行比较。ES最稳定地实现了良好的稳定性-可塑性权衡,而GA是最具可塑性的方法,但其遗忘程度高于ES。为解释这一现象,我们研究了每种方法解周围的回报景观。ES找到了最宽的邻域,即权重空间中扰动策略仍能解决任务的区域,且连续任务邻域之间的重叠大小与方法的稳定性-可塑性权衡相关。通过新颖性搜索奖励GA中的行为多样性,使群体更具可塑性,但以遗忘为代价。最后,RL中常见的可塑性损失症状并未转移到NE中。总体而言,这些结果确立了NE在持续任务变化下作为RL的竞争性替代方案,并表明在权重空间中进行扰动训练可能是一种更广泛的持续学习的有用机制。
英文摘要
Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms that search directly in weight space through mutation and selection over a population of neural networks. Across a wide array of environments and environmental changes, with policies ranging from a few hundred parameters to million-parameter networks, we compare evolution strategies (ES) and genetic algorithms (GAs) against state-of-the-art continual RL variants and population-based RL. ES most consistently achieves a good stability-plasticity trade-off, while the GA is the most plastic method but forgets more than ES. To explain this, we study the return landscape around each method's solutions. ES finds the widest neighborhoods, i.e.\ regions of weight space in which perturbed policies still solve the task, and the size of the overlap between the neighborhoods of consecutive tasks correlates with a method's stability-plasticity trade-off. Rewarding behavioral diversity in a GA through novelty search makes the population even more plastic, at the cost of forgetting. Finally, symptoms of plasticity loss commonly reported in RL do not transfer to NE. Overall, these results establish NE as a competitive alternative to RL under continual task changes, and suggest that training under perturbations in weight space may be a useful mechanism for continual learning more broadly.