PetriBench:对动态状态空间上LLM推理的基准测试
PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces
- ETH Zürich(苏黎世联邦理工学院)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有基准的局限,提出PetriBench,利用Petri网在动态状态空间上评估LLM推理,通过四个任务族和三级难度,发现准确率随难度下降,并揭示测试时计算与任务间的不同交互,为推理能力提供统一可扩展的评测框架。
AI中文摘要:
刻画LLM推理仍是一个开放挑战,因为许多现有基准隔离特定推理技能、依赖外部知识或扩展成本高昂。我们引入PetriBench,一个紧凑、完全自包含且可扩展的基准,用于评估LLM在动态状态空间上的推理,其使用Petri网——一种用于建模真实世界并发和分布式系统的成熟形式化方法。PetriBench将推理组织为四个任务族,其范围和时域各不相同,并通过增加结构复杂度生成易、中、难三个级别,且对照精确真值进行评估。在多样化的专有和开源权重模型中,准确率随难度增加而一致下降,而更难实例暴露出日益不同的任务特定能力画像。额外分析表明,测试时计算可提升性能,但不同推理任务间的交互方式不同,且程序化生成随结构复杂度产生平滑扩展。综上,这些结果表明PetriBench为探究LLM推理的优势、局限和扩展行为提供了一个统一且可扩展的设置。
英文摘要:
Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a compact, fully self-contained, and scalable benchmark for evaluating LLM reasoning over dynamic state spaces using Petri nets, a mature formalism for modeling real-world concurrent and distributed systems. PetriBench organizes reasoning into four task families varying by scope and temporal horizon, with Easy, Medium, and Hard levels generated by increasing structural complexity and evaluated against exact ground truth. Across a diverse set of proprietary and open-weight models, accuracy decreases consistently with difficulty, while harder instances expose increasingly distinct task-specific capability profiles. Additional analyses show that test-time compute improves performance but interacts differently with different reasoning tasks, and that procedural generation yields smooth scaling with structural complexity. Together, these results show that PetriBench provides a unified and extensible setting for probing the strengths, limits, and scaling behavior of LLM reasoning.