arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VeriSkill:一种用于程序验证技能的自进化框架

VeriSkill: A Self-Evolution Framework for Program Verification Skills

Changguo Jia, Tianqi Zhao, Zhiyou Xiao, Weiming Zhang, Minghui Zhou

arXiv 2607.27733首次发表:更新:

AI 中文总结

本文提出专为程序验证设计的自进化框架VeriSkill,可将验证失败归因于技能缺陷并迭代优化技能,实验显示其在多种场景下均优于基线方法。

AI 中文摘要

用大语言模型(LLM)智能体实现程序验证自动化,需要生成规约、注解、辅助引理及工具调用,所有这些都依赖可复用的技能。技能自进化是一种自然的解决方案,即从轨迹中提炼技能并通过反馈优化技能。然而,现有的进化方法在程序验证任务中表现不佳,因为它们无法可靠识别技能特定的失败,也无法从不透明的验证器反馈中提取可操作信号。本文提出专为程序验证设计的自进化框架VeriSkill,它将验证失败归因于技能缺陷,提炼诊断特征形成可复用的经验教训,并迭代优化候选技能,仅接受那些能提升验证性能同时保留程序语义的修订。实验表明,在多种验证工具、智能体框架及LLM后端上,VeriSkill始终优于所有基线方法。

英文摘要

Automating program verification with LLM agents requires generating specifications, annotations, auxiliary lemmas, and tool invocations, all of which depend on reusable skills. A natural remedy is skill self-evolution: distilling skills from trajectories and refining them through feedback. However, existing evolution methods struggle with program verification tasks because they cannot reliably identify skill-specific failures or extract actionable signals from opaque verifier feedback. In this paper, we propose VeriSkill, a self-evolution framework built for program verification. It attributes verification failures to skill deficiencies, distills diagnostic signatures into reusable lessons, and iteratively refines candidate skills, admitting only revisions that improve verification performance while preserving program semantics. Experiments show that VeriSkill consistently outperforms all baselines across multiple verification tools, agent frameworks, and LLM backends.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑