用于演化大规模基质的张量加速急切多分辨率网格
Tensor-Accelerated Eager Multi-Resolution Grids for Evolving Large-Scale Substrates
浏览论文内容
中文总结 AI 辅助
本文针对ES-HyperNEAT的四叉树难以张量化的问题,提出EMR-HyperNEAT,通过预先评估所有位置再过滤的方式实现并行,获得GPU加速并提升求解率。
中文摘要 AI 辅助
在神经演化中,间接编码从紧凑基因组生成神经网络连接,而非指定每个连接。ES-HyperNEAT通过检查CPPN输出模式自动发现隐藏节点的放置位置:它使用四叉树递归细分空间,在CPPN输出显示高方差的区域扩展。这种自适应方法无需手动指定基质即可发现网络拓扑,扩展了基于NEAT构建的固定网格HyperNEAT框架。然而,四叉树难以张量化:每个深度层级依赖父节点的方差,强制顺序评估;不同CPPN产生不同的细分模式,无法批量处理;且可变叶节点数不符合JAX的JIT编译对静态形状的要求。我们的前期工作证实,深度超过5时这些限制显现,尽管进行了批量优化,四叉树的JAX重新实现仅获得微小加速,这推动了本文提出的急切重构。我们提出EMR-HyperNEAT,它预先评估所有分辨率下的所有位置,再使用相同的方差准则进行过滤:将ES-HyperNEAT的“细分(若方差>θ)”变为“全部评估(eval_all());过滤(方差>θ)”。这会执行比必要更多的CPPN查询,但所有查询都相互独立,可跨核心和种群成员并行,将复杂度从O(4^D)降至P个并行核心下的O(4^D/P)。通过连接类型分类法,循环基质配置变得可行。实验部分验证,在深度5-7的XOR任务上,设备GPU加速达12-34倍,且在基准测试中求解率经验上更高。
英文摘要
In neuroevolution, indirect encoding generates neural network connectivity from a compact genome rather than specifying each connection. ES-HyperNEAT automatically discovers where to place hidden nodes by examining CPPN output patterns: it recursively subdivides space using a quadtree, expanding regions where CPPN outputs show high variance. This adaptive approach discovers network topology without manual substrate specification, extending the fixed-grid HyperNEAT framework built on NEAT. However, the quadtree resists tensorization. Each depth level depends on the parent's variance, forcing sequential evaluation. Different CPPNs produce different subdivision patterns, preventing batching. And variable leaf counts are incompatible with JAX's static shape requirement for JIT compilation. Our prior work confirmed these limits at depths exceeding 5, and a JAX reimplementation of the quadtree yielded only marginal speedup despite batched optimizations, motivating the eager reformulation presented here. We present EMR-HyperNEAT, which evaluates all positions at all resolutions up front, then filters using the same variance criterion: ES-HyperNEAT's subdivide_if(var > $θ$) becomes eval_all(); filter(var > $θ$). This performs more CPPN queries than necessary, but all queries become independent and parallelizable across both cores and population members, reducing complexity from \BigO($4^D$) to \BigO($4^D/P$) across $P$ parallel cores. Recurrent substrate configurations become feasible through a connection type taxonomy. The experiments section validates 12-34$\times$ on-device GPU speedup on XOR at depths 5-7, and empirically higher solve rates across benchmarks.