TLA$^{+}$-Bench:用于自然语言到TLA+规范生成的基于执行的基准测试和数据集
TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation
浏览论文内容
中文总结 AI 辅助
研究针对大语言模型编写TLA$^{+}$规范进展难测问题,提出TLA$^{+}$-Bench数据集和基准测试,基于执行评级。通过该基准测试发现评级选择影响正确率,存在正确性包络,且模型编写有效规范多但正确规范少,正确性随难度下降。
中文摘要 AI 辅助
大语言模型越来越多地根据自然语言描述编写TLA$^{+}$形式规范,但进展难以衡量:现有资源通过与参考的相似性或输出是否可解析来评级,均无法表明正确性。我们提出了TLA$^{+}$-Bench,一个基于执行进行评级的数据集和基准测试。每个黄金规范都附带一个配置,TLA$^{+}$模型检查器在整个可达状态空间上运行,以确定规范是否具有配置名称所对应的属性。该数据集包含来自13个公共存储库的403个经过模型检查的黄金规范和897个仅可解析的银色规范,包含了先前的TLA$^{+}$生成数据,并带有来自两个提供者的两种风格的四个模型编写描述,以及难度和类别标签。我们的主要发现是关于测量本身:一个精确的预言机给出的不是一个正确性数字,而是一个范围。仅改变早期基准测试未说明的评级选择,在一组固定的模型输出上,正确率变化了六倍,从10.0%到1.7%;加上接口供应选择,即告知模型配置的名称,范围扩大到十一倍,从18.7%到1.7%。我们将这个范围称为正确性包络,并测量其每个边界。其中的发现是稳定的。每个模型编写有效TLA$^{+}$的频率远高于正确TLA$^{+}$:最强的模型默认情况下正确率为16%,给出接口名称时为26%,开放模型最多为1%,并且正确性随着难度急剧下降。
英文摘要
Large language models increasingly write TLA$^{+}$ formal specifications from natural-language descriptions, but progress is hard to measure: existing resources grade by resemblance to a reference or by whether the output parses, neither of which shows correctness. We present TLA$^{+}$-Bench, a dataset and benchmark that grades by execution. Every gold specification ships a configuration the TLA$^{+}$ model checker runs over the full reachable state space, deciding exactly whether the specification holds the properties that configuration names. The dataset holds 403 model-checked gold and 897 parse-only silver specifications from 13 public repositories, subsumes prior TLA$^{+}$ generation data, and carries four model-written descriptions in two styles from two providers, with difficulty and category labels. Our main finding is about measurement itself: an exact oracle gives not one correctness number but a range. Varying only the grading choices earlier benchmarks leave unstated, on one fixed set of model outputs, the correct rate moves sixfold, from 10.0\% to 1.7\%; adding the interface-supply choice, where the model is told the configuration's names, widens the range to elevenfold, from 18.7\% to 1.7\%. We call this range the correctness envelope and measure each of its bounds. The findings inside it are stable. Every model writes valid TLA$^{+}$ far more often than correct TLA$^{+}$: the strongest is correct 16\% of the time by default and 26\% when given the interface names, open models at most 1\%, and correctness falls sharply with difficulty.
发表机构
- Loyola University Chicago(芝加哥洛约拉大学)
机构由 AI 辅助整理,请以论文原文为准。