arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ALPS:通过数学构造衡量大语言模型的有效创造力

ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction

Eric Xie, Wenqian Ye, Aidong Zhang

arXiv 2608.15979首次发表:更新:

AI 中文总结

本研究提出ALPS基准,通过要求生成可证正确的原创数学结构或证明其不存在的任务,评估大语言模型的有效创造力,测试发现现有推理模型在证明侧成功率14%、构造侧为0,多数实例未被解决并发布完整ALPS资源。

AI 中文摘要

大语言模型会生成被呈现为发现的输出,如新证明、猜想或分子。这类看似有创造力的输出是否真的原创且有效难以确定:开放式输出需要主观判断,可能复制训练中见过的内容,或任务过于简单无需创造力。我们提出ALPS(Austin-Law Proof-Synthesis,奥斯汀定律证明合成),这是一个用于衡量有效创造力的基准,其设计的任务要求生成既原创又可被证明正确的解决方案。每个实例为单个等式定律,经认证需构造满足该定律的无限数学结构,或证明不存在此类结构。提交内容由自动证明检查验证,无人工参与,且公开生成器可无限生成新实例,因此大语言模型不会在可能见过的问题上被评估。一组由8种领先自动证明器配置组成的组合,解决了4141个定律评估库中的2.2%,预算增加20倍仅多解决0.6%:障碍并非计算能力,而是缺乏能生成每个定律所需定制结构的方法。在固定协议下,我们测试的最强推理模型在证明侧的实例成功率为14%,但在构造侧无一成功。在我们测试的所有配置和预算下,库中剩余97.2%的实例均未被解决。我们完整发布ALPS:包括语料库、生成器和自动评判器。

英文摘要

Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears creative is truly original and effective is hard to establish: open-ended outputs require subjective judgment, the output may replicate something seen in training, or the task may be too simple to need creativity. We present ALPS (Austin-Law Proof-Synthesis), a benchmark that designs a task to measure valid creativity: producing a solution that is original and can be proven correct. Each instance is a single equational law, certified to require either the construction of an infinite mathematical structure satisfying the law, or a proof that no such structure exists. Submissions are verified by automated proof checking with no human involvement, and a public generator produces new instances without limit, so LLMs are never evaluated on problems they may have seen. A portfolio of eight configurations of leading automated provers resolves 2.2% of the 4,141-law evaluation pool, and a twentyfold budget increase adds 0.6%: the obstacle is not compute, but the absence of any method that produces the tailored structure each law requires. Under a fixed protocol, the strongest reasoning model we test succeeds in 14% of instances on the proof side, but none on the construction side. The remaining 97.2% of the pool is unresolved at every configuration and budget we test. We release ALPS in full: the corpus, the generator, and the automated judge.

Comments14 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑