arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SymbolicArena:符号回归中基准蒸馏与动态评估的统一基础设施

SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression

Ziwen Zhang, Xiju Wu, Yuheng Jing, Runxiang Wang, Boxiao Wang, Yifan Zang, Yifan Zhang, Yang Wang, Kai Li, Yifan Zhang, Huilin Xu, Jian Cheng

arXiv 2609.35113首次发表:更新:

发表机构

School of Artificial Intelligence, University of Chinese Academy of Sciences; CDL, Institute of Automation, Chinese Academy of Sciences; Nanjing University of Science and Technology; School of Future Technology, University of Chinese Academy of Sciences(中国科学院大学人工智能学院; 中国科学院自动化研究所CDL; 南京理工大学; 中国科学院大学未来技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SymbolicArena 提出统一基础设施,通过蒸馏 664 个任务为 Core50 基准,减少 92.5% 评估成本,并保持与完整任务集的一致性,同时揭示数值拟合与符号恢复间的显著差距。

AI 中文摘要

符号回归(SR)旨在从数据中寻找简洁且可解释的数学表达式,以用于科学方程发现。现有的符号回归基准在评估成本与基准有效性之间面临权衡。对大型任务池进行重复评估成本高昂,而紧凑型基准缺乏关于任务多样性和算法可区分性得以保留的系统性证据。SymbolicArena 提供了一个用于基准蒸馏和动态评估的统一基础设施。该框架将 664 个异构任务标准化,并配备可执行的真值表达式,同时将完整任务集蒸馏为 Core50,这是一个经过验证的包含 50 个任务的基准。蒸馏过程在明确的平衡约束下保留了任务覆盖范围和算法区分度。SymbolicArena 对异构符号回归算法应用统一的执行协议,并生成可比较的输出和搜索轨迹。多轴评估表征了数值质量、符号质量和搜索行为。Core50 将评估工作量减少了 92.5%,并保持与完整任务集评估的一致性。实验表明,SymbolicArena 的近似误差比替代选择器低 72.6% 至 86.7%,进一步支持其对完整任务集的保真度。评估揭示了当前符号回归方法在数值拟合与符号恢复之间存在显著差距,表明可靠的方程恢复仍然是一个开放的挑战。

英文摘要

Symbolic regression (SR) seeks concise and interpretable mathematical expressions from data for scientific equation discovery. Existing SR benchmarks face a tradeoff between evaluation cost and benchmark validity. Repeated evaluation of large task pools is expensive, and compact benchmarks lack systematic evidence of preserved task diversity and algorithm discriminability. SymbolicArena provides a unified infrastructure for benchmark distillation and dynamic evaluation. The framework standardizes 664 heterogeneous tasks with executable ground truth expressions and distills the Full Task Set into Core50, a validated benchmark of 50 tasks. The distillation process preserves task coverage and algorithm discrimination under explicit balance constraints. SymbolicArena applies a unified execution protocol to heterogeneous SR algorithms and produces comparable outputs and search trajectories. Multi Axis Evaluation characterizes numerical quality, symbolic quality, and search behavior. Core50 reduces evaluation workload by 92.5% and maintains agreement with Full Task Set evaluations. Experiments show that SymbolicArena achieves 72.6% to 86.7% lower approximation error than alternative selectors, further supporting its fidelity to the Full Task Set. Evaluation reveals a substantial gap between numerical fitting and symbolic recovery across current SR methods, suggesting that reliable equation recovery remains an open challenge.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑