arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越类型检查:迈向形式化规范生成的整体评估

Beyond Type-checking: Towards Holistic Evaluation of Formal Specification Generation

Srijith Nair, Aditya Vempaty, Jia Liu, Ashish Jagmohan

arXiv 2610.10604首次发表:更新:

发表机构

The Ohio State University; Emergence AI(俄亥俄州立大学; Emergence AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对形式化规范生成(SpecGen)缺乏整体评估的问题,构建含350个Lean任务的统一数据集与多维度评估框架,发现测量覆盖范围影响配置排名,且需区分输入输出相关指标以完善评估。

AI 中文摘要

在生成可验证代码时,需通过大语言模型(LLM)和智能体工作流将自然语言需求映射为机器可检查的代码。该流程的关键组件是规范生成(SpecGen),它会生成形式化契约,智能体可据此证明实现的正确性。证明生成能从定理证明器获取确定性反馈,但SpecGen缺乏对生成规范是否捕捉用户意图的明确检查,因此已验证的证明可能是基于歪曲预期行为的规范建立的正确性。本研究为实现SpecGen的整体评估,构建了一个统一数据集,包含350个现有Lean任务,其中189个来自VERINA、161个来自CLEVER,以及一个覆盖形式有效性、参考相似性与等价性、行为充分性的框架。我们区分了对所需输入的接受与对有效输出的接受、对无效输出的拒绝,同时明确了每个指标的证据范围。在四种SpecGen配置中,将参考相似性度量的广义树编辑距离(GTED)比较限制在32个可联合测量的VERINA任务后,VERINA配置的平均相似性排名从第二位降至第四位,凸显了测量覆盖范围的重要性。在一项人工控制实验中,某规范实现了100%的正测试召回率和负测试拒绝率,但仅接受0%的所需输入,这表明完美的后置条件分数可能忽略了不可用的输入契约,需对输入覆盖和输出约束分别提供反馈。

英文摘要

When generating verifiable code, natural language requirements are mapped to machine checked code using LLMs and agentic workflows. A crucial component of this pipeline is specification generation (SpecGen), which produces a formal contract against which an agent can prove implementation correctness. Proof generation can obtain deterministic feedback from a theorem prover, but SpecGen lacks a definitive check that a generated specification captures the user's intent. A checked proof can therefore establish correctness against a specification that misrepresents the intended behaviour. We take a step towards holistic SpecGen evaluation with a unified dataset assembled from $350$ existing Lean tasks, including $189$ from VERINA and $161$ from CLEVER, and a framework covering formal validity, reference similarity and equivalence, and behavioural adequacy. We distinguish acceptance of required inputs from acceptance of valid outputs and rejection of invalid outputs, while making each metric's evidence scope explicit. Across four SpecGen configurations, restricting the generalized tree edit distance (GTED) comparison, a reference similarity measure, to $32$ jointly measurable VERINA tasks changes the VERINA configuration's position from second to fourth in mean similarity, showing the importance of measurement coverage. In an authored control, a specification achieves $100\%$ positive test recall and negative test rejection while accepting $0\%$ of required inputs. This demonstrates that perfect postcondition scores can miss an unusable input contract, motivating separate feedback on input coverage and output constraints.

CommentsAccepted at NeurIPS 2026 Workshop on AI for Verifiable Coding (20 pages, 7 figures)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑