arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多样化以进行验证:当任务等效程序在可验证性上存在差异时

Diversifying to Verify: When Task-Equivalent Programs Differ in Verifiability

Shirley Yu, Ruben Martins

arXiv 2607.09366首次发表:更新:

发表机构

Stanford University; Carnegie Mellon University(斯坦福大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究任务等效程序的可验证性差异,提出基于大型语言模型的Diversify2Verify管道,通过推断契约、生成测试不同实现并修复注释来验证,构建基准测试,结果表明实现多样性有助于验证,提高了验证成功率。

AI 中文摘要

程序验证对软件正确性至关重要,但实际中生成完全验证的程序仍很困难。本文研究当多个生成的程序旨在满足相同任务级语义时,实现结构是否影响自动可验证性。我们提出了Diversify2Verify,这是一个基于大型语言模型的分阶段管道,用于Why3,它能推断特定表示的契约,生成并测试不同的递归和命令式数组/列表实现,并尝试通过有界验证器引导的注释修复进行验证。我们还构建了一个针对整数、数组和列表的73个任务的面向验证的基准,产生了292个实现变体。Diversify2Verify最初验证了96个工件,经过两次修复后验证了154个,将工件级验证从32.9%提高到52.7%。在任务级别,73个任务中有49个至少有一个变体得到验证,成功率为67.1%。这些结果表明任务等效的实现可验证性可能有很大差异,且实现多样性有助于找到便于验证的工件。

英文摘要

Program verification is crucial for software correctness, but producing fully verified programs remains difficult in practice. This paper studies whether implementation structure affects automated verifiability when multiple generated programs are intended to satisfy the same task-level semantics. We present Diversify2Verify, a staged LLM-based pipeline for Why3 that infers representation-specific contracts, generates and tests diverse recursive and imperative array/list implementations, and attempts verification with bounded verifier-guided annotation repair. We also construct a verification-oriented benchmark of 73 tasks over integers, arrays, and lists, yielding 292 implementation variants. Diversify2Verify verifies 96 artifacts initially and 154 after two repair passes, improving artifact-level verification from 32.9% to 52.7%. At the task level, at least one variant verifies for 49 of 73 tasks, a 67.1% success rate. These results show that task-equivalent implementations can differ substantially in verifiability and that implementation diversity helps find verification-friendly artifacts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑