发表机构
Shanghai AI Laboratory; The Hong Kong University of Science and Technology; University of California, Los Angeles(上海人工智能实验室; 香港科技大学; 加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出基于大型语言模型的残差感知科学方程发现方法RISR,通过残差编码器与双视图关系编码器优化公式发现,在LLM-SRBench任务上的ID、OOD准确率均优于基线,可提升数值方程恢复效果。
AI 中文摘要
符号回归结合了结构搜索与数值拟合,但聚合拟合分数无法描述剩余误差在不同输入上的变化情况。我们提出RISR,一种残差感知方法,利用这些误差模式指导公式发现并学习哪些修正值得拟合。残差编码器将对齐的输入、目标、当前预测及残差压缩为连续令牌,以条件化语言模型提出公式;后续优化阶段,双视图关系编码器利用加性和正则化乘性残差预测候选修正的拟合后效用。我们在LLM-SRBench的科学任务上评估RISR,RISR在1%和0.1%的逐点相对误差容限下分别达到63.57%和38.50%的ID准确率,对应的OOD准确率为56.07%和38.24%。RISR在使用相同主干的情况下优于已报告的基线,结果表明我们的残差感知方法可提升数值方程的恢复效果。
英文摘要
Symbolic regression combines structural search with numerical fitting, but aggregate fit scores do not describe how the remaining error varies across inputs. We introduce RISR, a residual-informed method that uses these error patterns to guide formula discovery and learn which corrections are worth fitting. A residual encoder compresses aligned inputs, targets, current predictions, and residuals into continuous tokens that condition a language model to propose formulas. For subsequent refinement, a dual-view relational encoder uses additive and regularized multiplicative residuals to predict the post-fit utility of candidate corrections. We evaluate RISR on scientific tasks from the LLM-SRBench. RISR achieves 63.57% and 38.50% ID accuracy at the 1% and 0.1% pointwise relative-error tolerances, respectively. The corresponding OOD accuracies are 56.07% and 38.24%. RISR outperforms the reported baselines using the same backbone. The results show that our residual-informed approach can improve numerical equation recovery.