arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FormalTCS:评估大型语言模型端到端前沿形式理论计算机科学研究的基准

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

Dingzirui Wang, Xuanliang Zhang, Keyan Xu, Qingfu Zhu, Wanxiang Che

arXiv 2608.20153首次发表:更新:

发表机构

Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FormalTCS基准评估发现,当前大型语言模型完成端到端前沿TCS研究的能力不足,形式化是核心瓶颈,且研究品味限制了自主TCS研究。

AI 中文摘要

大型语言模型(LLMs)在自动化理论计算机科学(TCS)研究中展现出日益增长的潜力,但现有基准仍远未贴合实际研究场景。我们推出\

英文摘要

Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings. We introduce \ourbenchmark, an expert-validated benchmark for evaluating LLMs on frontier, end-to-end TCS research. \ourbenchmark contains $143$ instances drawn from papers accepted to STOC, FOCS, SODA, and COLT in 2025-2026, preserving paper-specific definitions, assumptions, and proof dependencies, with expert-verified Lean formalizations and proofs. Evaluations of leading LLMs reveal that current models remain far from reliably completing the full research pipeline. In particular, autoformalization is the sharpest bottleneck: the best model achieves only $11.5$ on translating natural-language claims into formal theorem statements, compared with $28.6$ Pass@8 when proving human-provided formal statements. Building on \ourbenchmark, we further develop an automated TCS research framework that generates, formalizes, filters, and proves new claims. Of $64$ generated claims, only $6$ ultimately pass expert evaluation and proof verification, indicating that beyond formalization, limited research taste remains another major barrier to autonomous TCS research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑