arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SMILE:连接连续优化与离散符号恢复

SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery

Mansooreh Montazerin, Antonio Ortega, Ajitesh Srivastava

arXiv 2609.04639首次发表:更新:

发表机构

University of Southern California; Northeastern University(南加州大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SMILE是连接连续优化与离散符号恢复的混合框架,经SRBench评估,在高噪声下符号解率最高、鲁棒性强,且能快速恢复更简单的表达式,位于准确率与复杂度的帕累托前沿。

AI 中文摘要

符号回归(SR)从数据中发现闭式数学表达式,提供了黑盒模型之外的可解释性。现有方法在组合搜索空间中收敛缓慢,且缺乏利用数据中组合结构的机制。我们引入SMILE(Sine、Multiplication、Identity、Logarithm、Exponential),这是一种混合框架,通过三个阶段将基于连续梯度的优化与离散符号恢复统一起来:对数据进行结构分析以识别目标表达式的组合层次结构,连续优化以学习使用可解释激活函数编码目标表达式的网络参数,以及通过结构化剪枝、系数优化和舍入进行符号恢复。最后一阶段将学习到的网络提炼为具有精确符号常数的紧凑表达式。我们在SRBench上对SMILE进行评估,涉及真实值数据集和黑盒数据集, ablation研究验证了每个组件。SMILE在最大噪声水平下实现了最高的符号解率,在竞争方法大幅退化的情况下表现出强鲁棒性。它始终位于准确率与复杂度的帕累托前沿,以竞争方法所需时间的一小部分恢复出明显更简单的表达式。

英文摘要

Symbolic regression (SR) discovers closed-form mathematical expressions from data, offering interpretability beyond black-box models. Existing methods suffer from slow convergence in combinatorial search spaces and lack mechanisms to exploit compositional structure in the data. We introduce SMILE (Sine, Multiplication, Identity, Logarithm, Exponential), a hybrid framework that unifies continuous gradient-based optimization with discrete symbolic recovery through three stages: structural analysis of the data to identify the compositional hierarchy of the target expression, continuous optimization to learn parameters of a network that encodes the target expression using interpretable activations, and symbolic recovery through structured pruning, coefficient optimization, and rounding. This final stage distills the learned network into a compact expression with exact symbolic constants. We evaluate SMILE on SRBench across ground-truth and black-box datasets, with ablation studies validating each component. SMILE achieves the highest symbolic solution rate at the largest noise levels, demonstrating strong robustness where competing methods degrade substantially. It consistently lies on the Pareto front of accuracy versus complexity, recovering significantly simpler expressions in a fraction of the time required by the competing methods.

CommentsAccepted to NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑