arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23425cs.SEcs.AI

TLA$^{+}$-Bench:用于自然语言到TLA+规范生成的基于执行的基准测试和数据集

TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation

Arslan Bisharat, Eric Spencer, Brian Ortiz, Khushboo Bhadauria, Mujtaba Nazari, Beatriz Santos, Anisa Ramos, TaiNing Wang, George K. Thiruvathukal, Konstantin L… 展开作者

Arslan Bisharat, Eric Spencer, Brian Ortiz, Khushboo Bhadauria, Mujtaba Nazari, Beatriz Santos, Anisa Ramos, TaiNing Wang, George K. Thiruvathukal, Konstantin Läufer, Mohammed Abuhamad

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型编写TLA$^{+}$规范进展难测问题,提出TLA$^{+}$-Bench数据集和基准测试,基于执行评级。通过该基准测试发现评级选择影响正确率,存在正确性包络,且模型编写有效规范多但正确规范少,正确性随难度下降。

中文摘要 AI 辅助

大语言模型越来越多地根据自然语言描述编写TLA$^{+}$形式规范,但进展难以衡量:现有资源通过与参考的相似性或输出是否可解析来评级,均无法表明正确性。我们提出了TLA$^{+}$-Bench,一个基于执行进行评级的数据集和基准测试。每个黄金规范都附带一个配置,TLA$^{+}$模型检查器在整个可达状态空间上运行,以确定规范是否具有配置名称所对应的属性。该数据集包含来自13个公共存储库的403个经过模型检查的黄金规范和897个仅可解析的银色规范,包含了先前的TLA$^{+}$生成数据,并带有来自两个提供者的两种风格的四个模型编写描述,以及难度和类别标签。我们的主要发现是关于测量本身:一个精确的预言机给出的不是一个正确性数字,而是一个范围。仅改变早期基准测试未说明的评级选择,在一组固定的模型输出上,正确率变化了六倍,从10.0%到1.7%;加上接口供应选择,即告知模型配置的名称,范围扩大到十一倍,从18.7%到1.7%。我们将这个范围称为正确性包络,并测量其每个边界。其中的发现是稳定的。每个模型编写有效TLA$^{+}$的频率远高于正确TLA$^{+}$:最强的模型默认情况下正确率为16%,给出接口名称时为26%,开放模型最多为1%,并且正确性随着难度急剧下降。

英文摘要

Large language models increasingly write TLA$^{+}$ formal specifications from natural-language descriptions, but progress is hard to measure: existing resources grade by resemblance to a reference or by whether the output parses, neither of which shows correctness. We present TLA$^{+}$-Bench, a dataset and benchmark that grades by execution. Every gold specification ships a configuration the TLA$^{+}$ model checker runs over the full reachable state space, deciding exactly whether the specification holds the properties that configuration names. The dataset holds 403 model-checked gold and 897 parse-only silver specifications from 13 public repositories, subsumes prior TLA$^{+}$ generation data, and carries four model-written descriptions in two styles from two providers, with difficulty and category labels. Our main finding is about measurement itself: an exact oracle gives not one correctness number but a range. Varying only the grading choices earlier benchmarks leave unstated, on one fixed set of model outputs, the correct rate moves sixfold, from 10.0\% to 1.7\%; adding the interface-supply choice, where the model is told the configuration's names, widens the range to elevenfold, from 18.7\% to 1.7\%. We call this range the correctness envelope and measure each of its bounds. The findings inside it are stable. Every model writes valid TLA$^{+}$ far more often than correct TLA$^{+}$: the strongest is correct 16\% of the time by default and 26\% when given the interface names, open models at most 1\%, and correctness falls sharply with difficulty.

发表机构

  • Loyola University Chicago(芝加哥洛约拉大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑