发表机构
Harbin Institute of Technology(哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FormalTCS基准评估发现,当前大型语言模型完成端到端前沿TCS研究的能力不足,形式化是核心瓶颈,且研究品味限制了自主TCS研究。
AI 中文摘要
大型语言模型(LLMs)在自动化理论计算机科学(TCS)研究中展现出日益增长的潜力,但现有基准仍远未贴合实际研究场景。我们推出\
英文摘要
Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings. We introduce \ourbenchmark, an expert-validated benchmark for evaluating LLMs on frontier, end-to-end TCS research. \ourbenchmark contains $143$ instances drawn from papers accepted to STOC, FOCS, SODA, and COLT in 2025-2026, preserving paper-specific definitions, assumptions, and proof dependencies, with expert-verified Lean formalizations and proofs. Evaluations of leading LLMs reveal that current models remain far from reliably completing the full research pipeline. In particular, autoformalization is the sharpest bottleneck: the best model achieves only $11.5$ on translating natural-language claims into formal theorem statements, compared with $28.6$ Pass@8 when proving human-provided formal statements. Building on \ourbenchmark, we further develop an automated TCS research framework that generates, formalizes, filters, and proves new claims. Of $64$ generated claims, only $6$ ultimately pass expert evaluation and proof verification, indicating that beyond formalization, limited research taste remains another major barrier to autonomous TCS research.