AI 中文总结
该研究在LSR-Synth基准下,对比语言模型先验与常规算子搜索的效果,发现固定词汇表已覆盖多数任务,语言模型候选仅在词汇表覆盖受扰时贡献显著,且严格分布外评估不改变该关系。
AI 中文摘要
现有的科学方程发现基准大多由公共领域的知名方程组成,这使得难以判断模型是从数据中发现规律还是仅从训练语料库中回忆答案。LSR-Synth通过在已确立的科学机制中引入新型合成项,并对生成的任务进行新颖性、可解性和科学合理性过滤,缓解了这一问题。本文研究一个更狭窄的测量问题:这些任务能否进一步区分语言模型提供的科学先验与不访问任务语义的常规算子搜索?我们使用具有公开记录来源的固定词汇表构建无语义基线,并通过语义蒙蔽、库弱化和匹配算子族敲除评估候选覆盖的作用。在当前任务快照、搜索预算和评分协议下,固定词汇表已覆盖大多数任务,而语言模型生成的候选很少扩展可解实例集;仅当词汇表覆盖被选择性破坏时,其边际贡献才显著。严格分布外评估降低了所有方法的绝对成功率,但未改变该关系。这些发现既不否定LSR-Synth针对完整公式记忆的控制,也不意味着语言模型先验普遍无用,而是支持一个更有限的结论:大多数当前任务仍适合评估先前未见表达式的拟合与重组,但自身不足以识别固定搜索空间之外先验的贡献。
英文摘要
Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether a model is discovering laws from data or merely recalling answers from its training corpus. LSR-Synth mitigates this problem by introducing novel synthetic terms into established scientific mechanisms and filtering the resulting tasks for novelty, solvability, and scientific plausibility. This paper examines a narrower measurement question: can these tasks further distinguish scientific priors supplied by language models from conventional operator search that does not access task semantics? We construct a semantics-free baseline using a fixed vocabulary with publicly documented provenance, and assess the role of candidate coverage through semantic blinding, library weakening, and matched operator-family knockouts. Under the current task snapshot, search budget, and scoring protocol, the fixed vocabulary already covers most tasks, while language-model-generated candidates rarely expand the set of solvable instances. Their marginal contribution becomes substantial only when vocabulary coverage is selectively disrupted. Strict out-of-distribution evaluation lowers the absolute success rates of all methods but does not alter this relationship. These findings neither invalidate LSR-Synth's controls against memorization of complete formulas nor imply that language-model priors are generally unhelpful. Rather, they support a more limited conclusion: most current tasks remain suitable for evaluating the fitting and recombination of previously unseen expressions, but are insufficient on their own to identify contributions from priors beyond a fixed search space.