发表机构
University of Alberta; Filuta AI; University of British Columbia; Washington University in St. Louis(阿尔伯塔大学; Filuta AI; 不列颠哥伦比亚大学; 圣路易斯华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过塑造训练公式集并微调噪声数据,使Transformer在符号回归中合成物理公式,实现约十秒内生成且外推性能优于可比预算下的所有方法。
AI 中文摘要
找到一个紧凑的公式来拟合一组输入-输出对,并在未见过的输入上预测输出,是科学中的一个基本问题。符号回归自动化了对此类公式的搜索:基于搜索的方法直接探索可能的公式空间,而在合成数据上预训练的Transformer则能大幅更快地产生质量相当的公式。然而,现有的Transformer容易过拟合——它们找到的公式能很好拟合训练数据,但无法外推到训练中未见过的输入范围。我们通过塑造用于训练Transformer的公式集来解决这一问题,并表明由此产生的公式具有显著更好的外推能力。在带有噪声污染目标值的数据上对Transformer进行微调,进一步使合成的公式对观测中的噪声具有鲁棒性。在SRBench和LLM-SRBench上,我们的Transformer约在十秒内合成一个公式,并且在可比预算下外推性能优于所有被评估的方法。基于搜索的方法只有在获得一到三个数量级更多时间时才能超越我们的准确性。
英文摘要
Finding a compact formula that fits a set of input-output pairs and predicts outputs on unseen inputs is a fundamental problem in science. Symbolic regression automates the search for such formulae: search-based methods explore the space of possible formulae directly, while transformers pre-trained on synthetic data produce formulae of comparable quality substantially faster. Existing transformers, however, are prone to overfitting --- they find formulae that fit the training data well but do not extrapolate to input ranges unseen during training. We address this by shaping the set of formulae used to train a transformer, and show that the resulting formulae extrapolate substantially better. Fine-tuning the transformer on data with noise-corrupted target values further makes the synthesized formulae robust to noise in the observations. On SRBench and LLM-SRBench our transformer synthesizes a formula in about ten seconds and extrapolates better than all evaluated methods at a comparable budget. Search-based methods surpass our accuracy only when given one to three orders of magnitude more time.