无尽考试:从当今模型到超级智能的数学构造
The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
浏览论文内容
中文总结 AI 辅助
本文提出无尽考试基准,通过十四个参数化构造族和自动验证评分,衡量模型向超级智能的数学构造进展,并发布相关工具与数据。
中文摘要 AI 辅助
我们引入了“无尽考试”(Endless Exam),这是一个通过十四个参数化构造族来衡量从当今模型向人工超级智能取得数学进展的基准。每个提交的对象都会自动检查其有效性,并根据已发布的边界或构造基线获得相对质量分数,且改进上限不设限于1。这些构造族以开放数学问题为长期目标,并在更大参数下生成新实例,其中紧凑的证书使大型构造保持可验证性。在针对69个不同实例评估的八个模型中,连续质量分数能够区分性能,尽管没有评估系统超越已发布的边界。尺寸-质量曲线展示了随着问题规模增加,构造质量如何变化。我们发布了生成器、验证器、参考文献、模型响应和分析,以支持在人类边界之前和之后的持续测量。
英文摘要
We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be verified without listing every element. Across nine models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though none of the 30 published-frontier references is surpassed. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.
发表机构
- The University of Texas at Arlington(德克萨斯大学阿灵顿分校)
机构由 AI 辅助整理,请以论文原文为准。