AI 中文总结
研究符号回归遗传编程中零样本大语言模型生成亲本选择算子,在标准框架内对八个模型进行基准测试,比较不同模型生成算子在多个回归基准上的表现,发现Claude Sonnet~4.6和Gemini~3.1 Pro表现出色,证明零样本合成是可行方法。
AI 中文摘要
亲本选择对符号回归的遗传编程中的探索、利用和复杂性控制有显著影响。目前尚不清楚大语言模型能否在无迭代元进化的零样本设置中合成有效算子。本文在符号回归的标准遗传编程框架内,对八个大语言模型的亲本选择算子进行零样本合成基准测试。每个模型接收相同自然语言提示生成算子,在仅替换亲本选择算子的标准遗传编程框架中评估。在十二个回归基准上评估十个独立零样本算子,并与自动词法分析和锦标赛选择基线比较。Claude Sonnet~4.6和Gemini~3.1 Pro在训练和测试\(R^2\)上表现出色。基准中最强的Kimi~K2.5零样本合成算子在搜索有效性上超过基线。结果表明零样本大语言模型合成是生成有竞争力的遗传编程选择算子的可行方法。分析显示许多生成的算子使用语义指导选择,大语言模型可仅从任务描述产生非平凡搜索启发式。还研究了公共大语言模型排行榜排名与遗传编程性能的关系。如Humanity's Last Exam和SWE-bench Verified等广泛使用的基准与训练\(R^2\)强相关,与测试\(R^2\)的关系较弱且不明确。
英文摘要
Parent selection significantly affects exploration, exploitation, and complexity control in genetic programming (GP) for symbolic regression. It is unclear whether large language models (LLMs) can synthesize effective operators in a zero-shot setting without iterative meta-evolution. Here, zero-shot means that the model receives only the task description, with no reference operators or iterative feedback. In this work, we benchmark zero-shot synthesis of parent-selection operators across eight LLMs within a standard GP framework for symbolic regression. Each model receives the same natural-language prompt to generate a parent-selection operator, which is then evaluated in a standard GP framework with only the parent-selection operator replaced, while all other components and the evolutionary-search budget are held constant. For each LLM, ten independent zero-shot operators are evaluated on twelve OpenML regression benchmarks and compared against automatic lexicase and tournament selection baselines. Claude Sonnet~4.6 and Gemini~3.1 Pro stand out for consistently strong performance on both training and held-out test $R^2$. The strongest operator in our benchmark---a Kimi~K2.5 zero-shot synthesis---surpasses the automatic lexicase and tournament baselines in search effectiveness. These results suggest that zero-shot LLM synthesis is a viable approach to generating competitive GP selection operators. Analysis shows that many generated operators use semantics to guide selection, suggesting that LLMs can produce non-trivial search heuristics from the task description alone. We also examine the relationship between public LLM leaderboard rankings and GP performance. Widely used benchmarks, such as Humanity's Last Exam and SWE-bench Verified, strongly correlate with training $R^2$, while their relationship to held-out test $R^2$ is weaker and less clear.
CommentsAccepted at PPSN 2026