arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02628cs.LGcs.AI

符号回归中的深度分治方法

Deep Divide-and-Reduce in Symbolic Regression

Yusong Deng, Yanjie Li, Xin Ning, Lina Yu, Liping Zhang, Shu Wei, Mingzhu Wan, Min Wu, Weijun Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有符号回归方法的缺陷,提出DDRSR方法,拓宽表达式分解简化适用范围、规避暴力搜索,在相关任务中展现显著优势,明确适用场景与未来方向。

中文摘要 AI 辅助

符号回归(SR)是从数据中发现内在模式并以数学表达式表示的任务。当前用于符号回归的机器学习方法往往缺乏对支配这些表达式的内在数学和物理原理的深刻理解。尽管开创性的AI Feynman方法利用了数据背后的数学特性,但其表达式简化机制的适用范围较窄,在复杂方程上容易失效。此外,其底层机制严重依赖对子表达式的暴力搜索,极大限制了其实用性。通过严格的数学推导和证明,我们提出了符号回归中的深度分治方法(DDRSR)。DDRSR从根本上拓宽了表达式分解与简化的适用范围,规避了对子结构进行暴力搜索的需求,同时确保了更广泛的通用性和严格的理论正确性。实证评估表明,这些理论原理在表达式分解和数值回归任务中均展现出显著优势。最后,我们探讨了该范式的适用场景、固有局限以及未来研究的有前景方向。

英文摘要

Symbolic regression (SR) aims to discover underlying mathematical expressions from data while preserving interpretability. Most existing learning-based SR methods primarily optimize expressions from observations without explicitly exploiting their structural mathematical properties. AI Feynman introduced a complementary paradigm that leverages such properties to recursively decompose complex expressions, but its decomposition criteria cover only restricted structural forms and its treatment of nested composition can require brute-force search over candidate sub-expressions. Building on this paradigm, we propose Deep Divide-and-Reduce in Symbolic Regression (DDRSR), a mathematically grounded framework that systematically generalizes expression decomposition and variable reduction. DDRSR extends translational symmetry to coefficient- and exponent-interfered forms, enables variable separation under overlapping variables and additive constant offsets, and generalizes the identification of nested compositional structures. We further characterize an intrinsic non-identifiability limitation of decomposition when no effective variable separation is induced. Experiments across multiple symbolic regression algorithms and benchmark datasets show that DDRSR identifies a broader range of decomposable structures than AI Feynman and overall improves downstream regression accuracy and exact-expression recovery.

↑