arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新定位非线性:下游学习器如何重塑遗传编程必须进化的内容

Relocating Nonlinearity: How a Downstream Learner Reshapes What Genetic Programming Must Evolve

Nam H. Le

arXiv 2610.09347首次发表:更新:

发表机构

Vermont Complexity Center; University of Vermont(佛蒙特复杂性中心; 佛蒙特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过布尔域傅里叶度数实验,证明下游学习器决定遗传编程需构建的非线性程度,线性学习器迫使程序精确匹配目标度数,并指出学习器能力越强,程序可解读性越低。

AI 中文摘要

遗传编程最初被构想为一种进化解决方案的方式:程序本身就是答案,适应度则是其自身输出的误差。然而,大量后续工作转而将程序作为独立学习器的输入,使得适应度衡量的是学习器的输出而非程序本身的输出。选择何种学习器至关重要,但目前尚无理论解释这一选择依据。对于遗传编程而言,目标的难度取决于程序必须通过组合原语构建的非线性程度;学习器提供的任何非线性,程序都无需重复构建。因此,需要衡量的量是学习器能够接管目标中多少非线性,尽管在连续基准上只能进行估计。本文表明,程序仍需构建多少非线性由所连接的学习器决定,并且可以在程序本身上进行测量。转向布尔域后,目标的非线性恰好等于其傅里叶度数,我们在保持其他条件不变的情况下将该度数从一调节到六。在成功条件下,线性学习器迫使进化程序精确达到目标的度数,在每种设置下三十次运行中无一例外地覆盖一到六度;而树集成方法则以简单得多的程序成功,且成功率高出四倍。这为该领域带来了两点贡献:学习器可以根据目标的结构而非其难度声誉来选择,并且进化程序的度数是在任何运行过程中均可免费计算的诊断指标。我们的对照目标度数低于奇偶校验问题,却未从非线性学习器中获益,而可简化为简单统计量的目标则获益巨大。后者附带一个警告:进化仅在六十三个条件中的四个条件下内化了学习器提供的内容,并在四十四个条件下变得更加依赖学习器,因此学习器能力越强,模型中可在程序中解读的部分就越少。

英文摘要

Genetic programming was conceived as a way of evolving solutions: the program is the answer, and fitness is the error of its own output. A substantial line of work instead makes the program an input to a separate learner, so fitness measures the learner's output rather than the program's. Which learner to attach matters, with no account of what decides it. What makes a target hard for genetic programming is how much nonlinearity the program must build by composing primitives; whatever nonlinearity the learner supplies, the program need not. The quantity to measure is therefore how much of a target's nonlinearity a learner can take over, though on continuous benchmarks it can only be estimated. Here we show that how much the program must still build is decided by which learner is attached, and is measurable on the programs themselves. Moving to Boolean domains, where a target's nonlinearity is exactly its Fourier degree, we tune that degree from one to six with everything else fixed. Conditioned on success, a linear learner forces the evolved program to the target's degree exactly, at one through six without exception over thirty runs per setting, while tree ensembles succeed with far simpler programs and four times as often. This gives the field two things: a learner can be chosen from a target's structure instead of its reputation for difficulty, and the degree of the evolved program is a diagnostic free to compute during any run. Our control target has lower degree than the parity problems yet gains nothing from a nonlinear learner, while targets reducible to a simple statistic gain a great deal. The latter comes with a warning: evolution internalises what the learner supplies in only four of sixty-three conditions, and grows more dependent on it in forty-four, so the more capable the learner, the less of the model is legible in the program.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑