arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33594cs.CE

SymbolicLM:将语言模型训练为符号回归器

SymbolicLM: Training Language Models as Symbolic Regressors

Jun Yao, Yingfan Hua, Ruikun Li, Shixiang Tang, Bin Liu, Wanli Ouyang, Yan Lu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SymbolicLM,通过构建含超16万方程和18亿令牌的PhysSymbArena基准,以数值-符号和物理监督训练LLMs,并引入SymbolicSGA细化框架,显著提升符号回归的结构恢复能力。

中文摘要 AI 辅助

大型语言模型(LLMs)在科学推理方面已展现出令人瞩目的能力,然而科学发现最终要求直接从观测数据中推导出精确的定律,即符号回归(SR)。由于概率性文本生成与SR的精确结构要求之间存在差距,这对LLMs构成了挑战。现有方法依赖复杂的外部脚手架,这些方法计算成本高昂,且将符号推理与模型本身分离。为解决这一局限,我们提出通过专门的数值-符号和物理监督,直接赋予LLMs符号回归能力。我们引入了PhysSymbArena,一个大规模基准,包含超过160,000个方程和18亿个带有物理描述的数值-符号数据令牌,从而支持系统性的训练和评估。基于PhysSymbArena,我们开发了SymbolicLM,通过数学和物理监督增强LLMs的符号回归能力。在推理阶段,我们进一步引入SymbolicSGA,一个利用定量反馈迭代改进生成方程的细化框架。在多个符号回归基准上的实验表明,SymbolicLM在保持竞争力的数值拟合性能的同时,显著提高了结构恢复能力。这些结果表明,符号回归可以被明确地学习为LLMs的内在能力。

英文摘要

Large Language Models (LLMs) have shown promising capabilities in scientific reasoning, yet scientific discovery ultimately requires deriving precise laws directly from observational data, known as Symbolic Regression (SR). This poses a challenge for LLMs due to the gap between probabilistic text generation and the exact structural requirements of SR. Existing approaches rely on complex external scaffolds, which are computationally expensive and separate symbolic reasoning from the model itself. To address this limitation, we propose to directly equip LLMs with symbolic regression capabilities through dedicated numerical-symbolic and physical supervision. We introduce PhysSymbArena, a large-scale benchmark containing over 160,000 equations and 1.8B tokens of numerical-symbolic data with physical descriptions, enabling systematic training and evaluation. Based on PhysSymbArena, we develop SymbolicLM, which enhances the symbolic regression ability of LLMs through mathematical and physical supervision. During inference, we further introduce SymbolicSGA, a refinement framework that leverages quantitative feedback to iteratively improve generated equations. Experiments on multiple symbolic regression benchmarks show that SymbolicLM substantially improves structural recovery while maintaining competitive numerical fitting performance. These results demonstrate that symbolic regression can be explicitly learned as an intrinsic capability of LLMs.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

↑