arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpecAgent:通过智能体综合形式化程序规范赋能程序验证

SpecAgent: Empowering Program Verification with Agentic Synthesis of Formal Program Specifications

Lezhi Ma, Han Wang, Shangqing Liu, Jiawan Wang, Lei Bu

arXiv 2610.05132首次发表:更新:

发表机构

Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有LLM规范综合方法在复杂程序上易失效的问题,提出SpecAgent智能体框架,通过依赖感知规划、RAG、智能体修复与评判,在真实C程序上实现96.82%精确率与88.95%召回率,并成功解决604个验证目标。

AI 中文摘要

形式化规范对于演绎式程序验证至关重要,它为复杂软件的组合验证提供了语义抽象。然而,手动构建规范耗时费力,这促使了自动化综合的研究。尽管大型语言模型(LLM)近期取得了进展,但现有方法往往依赖于单向工作流和局部修复,这限制了它们在函数与循环之间存在复杂依赖关系的真实世界程序上的有效性。验证失败可能源于先前生成的规范,导致级联失败,而局部细化无法解决这些问题。为应对这些挑战,我们提出了SpecAgent,一个用于为真实世界C程序综合高质量ACSL规范的智能体框架。SpecAgent集成了四个组件:依赖感知规划以识别规范目标,检索增强生成(RAG)以提供相关规范模式和程序上下文,智能体修复以诊断缺陷并重新审视依赖的规范,以及智能体评判以评估和细化超出证明成功范围的语义强度。我们在规范综合和程序验证任务上评估了SpecAgent。在来自14个真实世界代码库的50个程序上,使用DeepSeek-V4的SpecAgent在综合正确且强健的规范方面达到了96.82%的精确率和88.95%的召回率,优于所有基线。在811个真实世界验证目标中,它成功解决了604个,也超越了现有基线。这些结果证明了SpecAgent在综合高质量规范和促进真实世界程序验证方面的有效性。

英文摘要

Formal specifications are essential for deductive program verification, providing semantic abstractions for compositional verification of complex software. However, manually constructing specifications is labor-intensive, motivating automated synthesis. Despite recent advances in large language models (LLMs), existing approaches often rely on forward-only workflows and localized repair, limiting their effectiveness on real-world programs with complex dependencies among functions and loops. Verification failures may stem from previously generated specifications, causing cascading failures that local refinement cannot resolve. To address these challenges, we present SpecAgent, an agentic framework for synthesizing high-quality ACSL specifications for real-world C programs. SpecAgent integrates four components: dependency-aware planning to identify specification targets, retrieval-augmented generation (RAG) to provide relevant specification patterns and program context, agentic repair to diagnose defects and revisit dependent specifications, and agentic critique to assess and refine semantic strength beyond proof success. We evaluate SpecAgent on specification synthesis and program verification tasks. On 50 programs from 14 real-world repositories, SpecAgent with DeepSeek-V4 achieves 96.82% precision and 88.95% recall in synthesizing correct and strong specifications, outperforming all baselines. On 811 real-world verification targets, it successfully discharges 604, also surpassing existing baselines. These results demonstrate SpecAgent's effectiveness in synthesizing high-quality specifications and facilitating real-world program verification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑