发表机构
Epoch AI(Epoch AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建OEIS Open基准,发现配备最少工具的语言模型在适中预算下可自主解决部分OEIS未解决数学猜想,访问大量数学文献或复杂智能体循环未提升其性能。
AI 中文摘要
我们构建了OEIS Open,这是一个基于OEIS中492个未解决数学猜想的基准,由Tsoukalas等人在Lean中形式化。此前这些猜想仅通过定制智能体尝试过,而我们的开源评估代码可让任何通用语言模型(LM)在该基准上运行,且能抵御LM的作弊尝试。我们发现,配备最少工具集的LM在每次尝试50美元的预算下,解决了其中147个猜想,在OEIS Open上的得分为30%。OEIS Open Lite是100个猜想的随机子集,用于更廉价的评估。当每次尝试预算为200美元时,当前最优LM在OEIS Open Lite上的得分为44%。让LM访问arXiv的47.6万篇数学文献,并未提升其在OEIS Open Lite上的性能,使用更复杂的智能体循环也未提升。本研究涵盖的猜想数学意义尚不明确,且多数此前可能很少受到关注。不过,我们的结果表明,LM能够以适中成本自主解决未解决的研究猜想。
英文摘要
We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Whereas these conjectures had previously been attempted only with a bespoke agent, our open-source evaluation code runs any generic language model (LM) against them, and is secure against LM cheating attempts. We find that LMs equipped with a minimal set of tools resolve 147 of these conjectures with a budget of \$50 per attempt, scoring 30% on OEIS Open. OEIS Open Lite is a random subset of 100 conjectures for cheaper evaluation. When evaluated with a budget of \$200 per attempt, the best current LM scores 44% on OEIS Open Lite. Giving LMs access to the mathematics literature via 476,000 papers from arXiv did not increase performance on OEIS Open Lite, and nor did using more sophisticated agent loops. The conjectures covered in this work are of uncertain mathematical significance, and most have likely received little previous attention. Nevertheless, our results show that LMs can resolve open research conjectures autonomously and at modest cost.
Comments26 pages, 6 figures. Code: https://github.com/epoch-research/LeanOpenProblems, results: https://github.com/epoch-research/LeanOpenProblems-results