发表机构
University of York(约克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Neuroevolution Arena嵌套评估协议,对比三种更新继承机制在两种架构上的表现,发现RL机制训练适应度更高,协议可分离训练工件与评估情境以暴露变异来源。
AI 中文摘要
竞争型人工生命系统在训练评估与生态评估下对训练好的控制器的排名可能存在差异。我们提出了Neuroevolution Arena,这是一个GPU加速的、由独立参数化神经网络单元构成的空间生态系统,以及一个带审计追踪的嵌套评估协议。我们将三种特定实现的更新与继承机制(EvoEvo、EvoRL和RLRL)与两种神经网络架构进行交叉组合,每种条件下开展三次独立训练运行,共50000代。18次运行中各保存一个精英控制器工件,进入对齐运行的冻结评估设计,包含198项计算任务。成对效应在每个对齐训练运行模块内平均了三个由种子定义的生态情境(两个允许合作、一个允许攻击);独立水平下每个条件的运行数n=3。启用强化学习(RL)的机制达到的记录训练适应度高于EvoEvo,而成对结果显示出架构依赖的多数模式和显著的工件依赖性。六方赢家随工件和情境变化,且预先指定的存活终点完全处于最低水平。我们贡献了一种嵌套协议,该协议将训练运行工件与评估情境分离,暴露而非掩盖它们不同的变异来源。
英文摘要
Competitive artificial-life systems can rank trained controllers differently under training and ecological evaluation. We present Neuroevolution Arena, a GPU-accelerated spatial ecology of independently parameterized neural-network cells, and an audit-tracked nested evaluation protocol. Three implementation-specific update-and-inheritance regimes (EvoEvo, EvoRL, and RLRL) are crossed with two neural architectures for 50,000 generations in three independent training runs per condition. One saved elite-controller artifact from each of the 18 runs enters an aligned-run frozen-evaluation design comprising 198 computational jobs. Pairwise effects average three seed-defined ecological contexts (two cooperation-permitting and one attack-permitting) within each aligned training-run block; the independent level remains n = 3 runs per condition. RL-enabled regimes attain higher recorded training fitness than EvoEvo, whereas pairwise outcomes show architecture-conditioned majority patterns and substantial artifact dependence. Six-way winners vary across artifacts and contexts, and the prespecified survival endpoint has a complete floor. We contribute a nested protocol that separates training-run artifacts from evaluation contexts and exposes, rather than conceals, their different sources of variation.
Comments22 pages, 2 figures, 8 tables. Preprint