arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09264cs.CLcs.LO

StochBench:Lean 中随机过程的领域特定基准

StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean

  • Case Western Reserve University(凯斯西储大学)

机构由 AI 辅助整理,请以论文原文为准。

Idan Davidovich, Debargha Ganguly, Vikash Singh, Vipin Chaudhary

AI总结:

StochBench是Lean 4中450个研究生级随机过程问题的领域特定基准,覆盖多种过程,Opus 4.8智能体证明率达34.9%,弥补了现有基准的领域空白。

AI中文摘要:

大型语言模型用于形式定理证明的领先基准是从竞赛数学(如IMO和Putnam)中提取的小型集合,这些基准不能很好地代表特定领域的应用。我们引入了StochBench,一个Lean 4基准,包含450个不同抽象层次的研究生级随机过程问题,每个问题都配有自然语言来源。针对Mathlib中代表性不足的领域,它涵盖了有限和可数马尔可夫链、更新过程、随机游走、鞅、停时、排队、布朗运动、随机微积分、弱收敛以及泊松和连续时间马尔可夫过程。我们基于Opus 4.8的智能体在每题15分钟的限制下达到了34.9%的证明率(157/450)。StochBench更好地代表了领域特定的应用数学,同时对高级证明器仍具挑战性。

英文摘要:

Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at varying abstraction levels, each paired with its natural-language source. Addressing a field underrepresented in Mathlib, it covers finite and countable Markov chains, renewal processes, random walks, martingales, stopping times, queues, Brownian motion, stochastic calculus, weak convergence, and Poisson and continuous-time Markov processes. Our Opus 4.8-based agent achieves a 34.9% proof rate (157/450) under a 15-minute per-problem limit. StochBench better represents domain-specific applied mathematics while remaining challenging for advanced provers.

↑