arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36187cs.LG

EvoMO-SR:基于子结构引导的多目标LLM符号表达式进化

EvoMO-SR: Multiobjective LLM-based Evolution of Symbolic Expressions with substructure guidance

Cristina Rossetti, Anna V. Kononova, Thomas Bäck, Fei Liu, Niki Van Stein

首次发表
浏览论文内容

中文总结 AI 辅助

提出EvoMO-SR框架,利用LLM生成方程骨架并外部拟合系数,通过多目标选择和子结构引导提升符号回归精度与结构恢复,在LSR-Synth多数设置中取得最佳性能。

中文摘要 AI 辅助

符号回归(SR)是一种用于科学发现的数据驱动方法,旨在数据中搜索可解释的分析关系。近年来,大型语言模型(LLMs)也对科学发现产生了显著影响,使得该过程的各个阶段得以自动化。基于这些原因,利用LLMs所蕴含的科学知识和编程能力来解决SR任务的可能性已经出现,与传统方法相比显示出有前景的性能。我们提出了EvoMO-SR,一种新颖的LLM驱动的SR框架,其中LLM生成方程骨架,其系数由外部优化器分别拟合。该框架包括一个多目标生存选择,通过平衡精度和复杂度来控制膨胀,以及一个子结构引导机制,用候选可重用构建块来变异表达式。EvoMO-SR在LSR-Synth的八个域内和域外设置中的七个中达到了最佳精度,使用的是小型LLM模型,即Llama-3.1-8B-Instruct。我们还通过基于规范化子树重叠和术语匹配的两个符号精度指标评估了结构恢复,表明我们的方法有更高的概率恢复高精度的符号结构。

英文摘要

Symbolic Regression (SR) is a data-driven method for scientific discovery which searches for interpretable analytical relationships within data. Recently, Large Language Models (LLMs) have also had a significant impact on scientific discovery, enabling the automation of various stages of the process. For these reasons, the possibility of harnessing the embedded scientific knowledge and programming capabilities of LLMs to solve SR tasks has emerged, showing promising performance compared with traditional methods. We propose EvoMO-SR, a novel LLM-driven SR framework in which the LLM generates equation skeletons, with their coefficients fitted separately by an external optimizer. The framework includes a multi-objective survival selection which controls bloating by balancing accuracy and complexity, and a substructure guidance mechanism which mutates expressions with candidate reusable building blocks. EvoMO-SR achieves the best accuracy in seven of the eight in-domain and out-of-domain settings for LSR-Synth, using a small LLM model, i.e., Llama-3.1-8B-Instruct. We also evaluated structural recovery through two symbolic accuracy metrics based on canonicalized subtree overlap and term matching, showing that our method has a greater probability of recovering highly accurate symbolic structures.

发表机构

  • Polytechnic Institute of Turin(都灵理工大学)
  • Leiden University(莱顿大学)
  • University of Zurich(苏黎世大学)
  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑