遗传编程符号回归中缓存策略的内存-运行时权衡分析
Analysis of Memory-Runtime Trade-offs in Caching Strategies for Genetic Programming Symbolic Regression
AI总结:
本文针对遗传编程符号回归的缓存机制展开全面分析,结合实证研究给出缓存配置指南,发现不同缓存策略在内存-运行时权衡上存在差异,轻量级策略可有效降低适应度评估时间。
AI中文摘要:
遗传编程符号回归(GPSR)通过进化过程生成数学表达式以建模输入输出关系,其核心挑战在于需重复评估完整表达式或其子表达式,这会增加计算运行时间。为解决该低效问题,研究中采用缓存机制减少冗余计算,但现有研究多仅使用单一缓存策略,对其相对性能或内存-运行时权衡的分析有限。本文针对GPSR在合成数据集和真实世界数据集上的缓存机制展开全面分析,同时开展了无限大缓存下键值使用频率的实证研究,为最优缓存大小的确定提供了依据,还给出了基于计算和内存约束配置缓存策略的可操作指南。研究发现,复杂缓存机制需至少达到特定缓存大小才能实现计算时间的减少;而轻量级缓存策略,如最近最少使用(LRU)策略,尤其是先进先出(FIFO)策略,可显著降低适应度评估的计算时间,适应度评估是整体运行时间的重要组成部分。
英文摘要:
Genetic Programming Symbolic Regression (GPSR) generates mathematical expressions to model input-output relationships using an evolutionary process. A significant challenge in GPSR lies in the repeated evaluation of entire expressions or their sub-expression, which inflates computational runtime. To address this inefficiency, caching mechanisms have been employed to reduce redundant computations. However, prior studies predominantly employ a single caching strategy, offering limited insights into their comparative performance or memory-runtime trade-offs. In this paper, we present a comprehensive analysis of caching mechanisms for GPSR on synthetic and real-world datasets. We also include an empirical study of key-value usage frequencies under an infinitely large cache, offering insights into optimal cache sizing. Furthermore, we provide actionable guidelines for configuring caching strategies based on computational and memory constraints. Our findings indicate that complex caching mechanisms necessitate a minimum cache size to achieve computational time reductions. Conversely, lightweight caching strategies, such as Least Recently Used (LRU) and, notably, First-In-First-Out (FIFO), can significantly decrease computation time for fitness evaluations, which are a substantial component of the overall runtime.