发表机构
The Hong Kong Polytechnic University; University of Washington; Amazon Web Services; Jiangxi University of Finance and Economics; EEO Education Technology; Chongqing Medical University; Eastern Institute of Technology(香港理工大学; 华盛顿大学; 亚马逊网络服务; 江西财经大学; EEO教育科技; 重庆医科大学; 东方理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出驻留GPU的批量Levenberg–Marquardt求解器,集成到EvoGP中,在符号回归任务中大幅提升常数优化效率,恢复控制方程的问题数量远超原生EvoGP。
AI 中文摘要
常数优化用于细化符号回归中树型遗传编程候选表达式的数值系数,但其每代计算成本过高,导致现代GPU加速框架要么省略该步骤,要么将其限制为轻量形式。本文提出一种驻留GPU的批量Levenberg–Marquardt求解器,该求解器每轮迭代使用固定数量的全局CUDA启动,对结构异构的表达式树种群进行常数优化。反向模式自动微分通过一次反向传播组装每棵树的雅可比矩阵,使每轮迭代的主要计算成本与每棵树的常数数量无关;双精度交付保障确保返回的常数不会比初始值更差。在早期种群阶段,该求解器在NVIDIA A100上可维持每秒多达5.1×10^5棵树的处理速度;在GPU饱和基准配置下,其吞吐量约为64核EPYC 7763上运行的Operon的9.9倍,同时匹配fp64参考质量。该求解器以进程内方式集成到EvoGP中,在18个构造问题中,其能恢复控制方程的问题数量为10个,而原生EvoGP仅为0个。
英文摘要
Constant optimization refines the numerical coefficients of candidate expressions in tree-based genetic programming for symbolic regression. But its per-generation cost has led modern GPU-accelerated frameworks to omit it or restrict it to lightweight forms. We present a GPU-resident, batched Levenberg--Marquardt solver that optimizes constants across a structurally heterogeneous population of expression trees using a fixed number of population-wide CUDA launches per iteration. Reverse-mode automatic differentiation assembles the per-tree Jacobian in one backward sweep, making the dominant per-iteration cost independent of the number of constants per tree, and a double-precision delivery guard guarantees that returned constants are never worse than their initial values. On early-generation populations, the solver sustains up to $5.1{\times}10^{5}$ trees per second on an NVIDIA A100; at a GPU-saturated benchmark configuration it delivers roughly $9.9{\times}$ the throughput of Operon running on a 64-core EPYC 7763, while matching fp64-reference quality. Integrated in-process into EvoGP, the solver enables end-to-end search to recover governing equations on $10$ of $18$ constructed problems versus 0 for stock EvoGP. Our code is at https://github.com/TensorConv/CuSR.
CommentsAccepted at the 30th Annual IEEE High Performance Extreme Computing Conference (HPEC 2026), 14-18 September 2026. To appear in IEEE Xplore