arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

科学方程发现的测试时缩放

Test-Time Scaling for Scientific Equation Discovery

Haowei Lin, Hubert Lim, Xiangyu Wang, Letian Huang, Di He

arXiv 2608.28660首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文将测试时缩放(TTS)应用于开放场景的自动方程发现,建模迭代搜索过程并在固定预算下对比控制器,发现搜索宽度是主导分配参数,为缩放LLM基方程发现提供了核心思路。

AI 中文摘要

测试时缩放(TTS)通过分配额外的测试时计算资源提升语言模型的推理能力,但现有研究主要聚焦于数学、编码等封闭任务。本文研究TTS在自动方程发现这一开放场景中的应用,该场景下模型需在候选方程中搜索并依赖观测数据获取反馈。我们将大语言模型(LLM)驱动的方程发现建模为迭代搜索过程,在统一的计算分配视角下整合了N-best、序列优化、树搜索及进化式方法。为隔离计算分配对提示工程及其他启发式方法的影响,我们在固定预算下对比了极简并行控制器。在LLM-SRBench方程发现任务上,我们发现搜索宽度是主导的分配参数:在我们的搜索范围内,最优宽度通常随计算预算增加而提升,而种群-分支划分及控制器选择的影响较小;合适的宽度选择还能通过增加并行度提升 wall-clock 效率。这些结果表明,给定一个具备信息性的验证器,控制探索与利用是缩放LLM基方程发现的核心。

英文摘要

Test-time scaling (TTS) improves language model reasoning by allocating additional test-time compute, but prior work mainly studies closed-ended tasks such as math and coding. We study TTS for automated equation discovery, an open-ended setting where models search over candidate equations and rely on observed datapoints for feedback. We formulate LLM-driven equation discovery as an iterative search process that unifies Best-of-N, sequential refinement, tree search, and evolution-style methods under a common compute-allocation view. To isolate allocation effects from prompt engineering and other heuristics, we compare minimal parallel controllers under fixed budgets. On LLM-SRBench equation-discovery tasks, we find that search width is the dominant allocation parameter: the best width in our sweep generally increases with the compute budget, while the population--branching split and controller choice matter less. Appropriate width selection also improves wall-clock efficiency by increasing parallelism. These results suggest that, given an informative verifier, controlling exploration and exploitation is central to scaling LLM-based equation discovery.

Journal refEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑