arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2502.05344cs.SEcs.AI

RAG-Verus:使用检索增强生成的大语言模型仓库级程序验证

RAG-Verus: Repository-Level Program Verification with LLMs using Retrieval Augmented Generation

  • University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

Sicheng Zhong, Jiading Zhu, Yifang Tian, Xujie Si

更新

AI总结:

针对现有函数中心方法忽略的跨模块依赖与全局上下文挑战,提出RagVerus框架,结合检索增强生成与上下文感知提示实现多模块仓库自动证明合成,在新RepoVBench基准获27%相对提升,受限预算下将现有基准证明通过率提高三倍。

AI中文摘要:

将自动化形式化验证扩展到实际项目需解决跨模块依赖和全局上下文问题,这是现有以函数为中心的方法忽略的挑战。我们提出RagVerus框架,将检索增强生成与上下文感知提示相结合,实现多模块仓库的自动证明合成,在新RepoVBench基准(首个Verus仓库级数据集,含383个证明完成任务)上相对提升27%;在受限语言模型预算下,RagVerus将现有基准的证明通过率提高三倍,展现出可扩展且样本高效的验证能力。

英文摘要:

Scaling automated formal verification to real-world projects requires resolving cross-module dependencies and global contexts, which are challenges overlooked by existing function-centric methods. We introduce RagVerus, a framework that synergizes retrieval-augmented generation with context-aware prompting to automate proof synthesis for multi-module repositories, achieving a 27% relative improvement on our novel RepoVBench benchmark -- the first repository-level dataset for Verus with 383 proof completion tasks. RagVerus triples proof pass rates on existing benchmarks under constrained language model budgets, demonstrating a scalable and sample-efficient verification.

↑