arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TCSAlgBench:面向研究级理论计算机科学的自动化证明基准测试

TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science

Chutong Yang, Xiyuan Zhang, Yu Huang, Boran Han, Soonho Kong, Shuai Zhang, Vihang Prakash Patil, Zhen Han, Michael Bohlke-Schneider, Bernie Wang

arXiv 2609.35606首次发表:更新:

发表机构

The University of Texas at Austin; Amazon; University of Pennsylvania(德克萨斯大学奥斯汀分校; 亚马逊; 宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TCSAlgBench是一个包含398个定理级挑战的基准测试,用于评估大型语言模型在研究级理论计算机科学中的自动化证明能力,实验显示智能体规划达到最高25.4%的验证者接受覆盖率。

AI 中文摘要

大型语言模型在竞赛数学中表现强劲,但其研究级推理能力仍难以被系统性地评估。理论计算机科学(TCS)将算法设计与明确的保证和基本限制联系起来,为评估模型能否用人类可检查的论证来证明计算改进提供了环境。我们引入了TCSAlgBench,一个用于自然语言证明发现的基准测试和可复用流程,包含来自138篇STOC和COLT 2026论文的398个定理级挑战。专家设计的规则完善了论文特定上下文,保留了计算假设和定量保证,并在发现算法是任务的一部分时扣留构造。对于每个任务,证明系统接收定理陈述并访问引用的先前工作。该流程支持从新发布论文中生成新鲜的、版本化的挑战批次。我们在直接推理和证明者-验证者讨论下评估了来自四个家族的十种模型配置,并在匹配的模型调用机会下比较了四种智能体工作流。所有评估均使用完整基准测试。在模型比较中,GPT-5.6 Sol max在10轮讨论后实现了最高的五次运行验证者接受覆盖率,为23.6%。讨论和重复采样提高了覆盖率。在使用GPT-5.5 xhigh的单独智能体比较中,分解优于讨论,而智能体规划实现了最高的五次运行验证者接受覆盖率,为25.4%。TCSAlgBench提供了一个可刷新的测试平台,用于衡量模型推理的进展并研究智能体工作流如何支持研究级证明发现。

英文摘要

Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify computational improvements with arguments humans can inspect. We introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery, comprising 398 theorem-level challenges from 138 STOC and COLT 2026 papers. Expert-designed rules complete paper-specific context, preserve computational assumptions and quantitative guarantees, and withhold constructions when discovering an algorithm is part of the task. For each task, prover systems receive theorem statements and access to cited prior work. The pipeline supports fresh, versioned challenge batches from newly released papers. We evaluate ten model configurations from four families under direct inference and prover-verifier discussion, and compare four agent workflows under matched model-call opportunities. All evaluations use the full benchmark. In the model comparison, GPT-5.6 Sol max achieves the highest five-run verifier-accepted coverage at 23.6% after 10-round discussion. Discussion and repeated sampling improve coverage. In the separate agent comparison using GPT-5.5 xhigh, decomposition improves coverage over discussion, and agentic planning achieves the highest five-run verifier-accepted coverage at 25.4%. TCSAlgBench provides a refreshable testbed for measuring progress in model reasoning and studying how agent workflows support research-level proof discovery.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑