发表机构
University of Wisconsin-Madison; University of Illinois at Urbana-Champaign(威斯康星大学麦迪逊分校; 伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何实现自然语言断言的自动形式化,提出Monty框架,利用新颖度量和有效性分数过滤形式化,经实验验证,该技术在断言生成任务中比单纯用大语言模型更可靠,平均精度提高多达20分。
AI 中文摘要
形式合同对软件测试和验证至关重要,但编写它们既费力又容易出错。大语言模型为自动形式化提供了一条有前景的途径:从自然语言规范中合成可执行断言,弥合非正式开发者意图与形式可执行规范之间的差距。我们提出了Monty:一个用于断言的自动形式化框架,应对断言有效性期望和自然语言歧义的挑战。我们的技术基于使用新颖的一致性分数度量和通过针对形式化断言测试代码获得的有效性分数来过滤形式化。我们在从22个类似集合的Java类派生的541个断言生成任务上评估了我们的方法,结果表明我们的技术比单纯使用大语言模型翻译断言更可靠地生成基本事实(平均精度提高多达20分)。
英文摘要
Formal contracts are essential for software testing and verification, yet writing them remains labor-intensive and error-prone. LLMs offer a promising path toward autoformalization: synthesizing executable assertions from natural-language specifications and thereby bridging the gap between informal developer intent and formal executable specifications. We present Monty: an autoformalization framework for assertions that tackles the challenges of expectations of validity of assertions and ambiguity in natural-language. Our techniques are based on filtering formalizations using a novel conformance score metric and validity scores obtained from testing the code against formalized assertions. We evaluate our approach on 541 assertion-generation tasks derived from 22 collection-like Java classes, and show that our technique produces the ground truth more reliably (improving upto 20 points in precision on average) than when using LLMs naively to translate assertions.