arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25291cs.LGcs.AI

InsightSR:通过并行语义与结构大语言模型引导优化符号回归搜索空间

InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance

Yating Ling, Wenjing Cun, Zhitang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

InsightSR是嵌入PySR遗传编程引擎的框架,通过LLM的语义与结构引导优化符号回归搜索空间,在多个基准测试中性能优于现有方法且泛化能力强。

中文摘要 AI 辅助

符号回归(SR)旨在从观测数据中发现简洁的数学规律,但传统方法常受困于具有物理意义表达式的庞大组合搜索空间。我们提出InsightSR,这一框架将大语言模型(LLMs)作为引导层嵌入PySR遗传编程引擎。InsightSR不依赖LLMs直接生成表达式,而是通过两条互补路径逐步变换搜索空间:语义种子路径提出维度一致的函数骨架,结构特征路径推荐非线性特征变换。这些变换在迭代中累积,拓宽输入空间,将符号搜索从基于原始变量构建深层表达式树,转变为基于丰富的、语义信息丰富的特征集组装浅层树。生成后反馈循环评估候选,按经验效用对特征分类,并优化下一轮引导,将发现过程从无约束生成转变为迭代、自校正的优化。在三个基准测试中,InsightSR在Feynman基准上实现95%的精确恢复率,在LLM-SRBench的LSR-Transform任务上达到80.18%的准确率,显著优于最先进的遗传编程和神经符号方法,同时在真实世界数据集上保持强泛化能力。

英文摘要

Symbolic regression (SR) seeks to discover parsimonious mathematical laws from observational data, yet conventional approaches often struggle with the vast combinatorial search space of physically meaningful expressions. We present InsightSR, a framework that embeds Large Language Models (LLMs) as a guiding layer around the PySR genetic programming engine. Rather than relying on LLMs to generate expressions directly, InsightSR uses LLMs to progressively transform the search space itself through two complementary pathways: a Semantic Seed Pathway that proposes dimensionally consistent functional skeletons, and a Structural Feature Pathway that recommends nonlinear feature transformations. These transformations accumulate over iterations, broadening the input space and shifting the symbolic search from constructing deep expression trees over raw variables to assembling shallow trees over a rich, semantically informed feature set. A post-generation feedback loop evaluates candidates, categorizes features by their empirical utility, and refines the guidance for the next iteration, transforming the discovery process from open-ended generation into iterative, self-correcting refinement. Across three benchmarks, InsightSR achieves a 95% exact recovery rate on the Feynman benchmark and 80.18% accuracy on the LLM-SRBench LSR-Transform task, substantially outperforming state-of-the-art genetic programming and neural-symbolic methods while maintaining strong out-of-distribution generalization on real-world datasets.

↑