arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Schwarz:感知求解器的智能体程序验证工具

Schwarz: Solver-Aware Agentic Program Verification

Jingyu Ke, Ling-I Wu, Guoqiang Li

arXiv 2608.30803首次发表:更新:

AI 中文总结

本文提出感知求解器的智能体程序验证工具Schwarz,通过本地化修复任务等方法,在两类基准任务上分别实现95.2%、91.5%的解决率,验证了其有效性与可扩展性。

AI 中文摘要

智能体验证系统常生成看似合理的源代码级规范,但合理性不足:验证器需将这些规范转化为求解器可证明的SMT约束。当此步骤失败时,当前大语言模型驱动的循环通常仅暴露粗略的验证器错误、超时或未知求解器结果,无法判断是规范错误、缺少辅助引理、证明上下文包含无关事实,还是约束需要不同的理论视图。本文提出Schwarz,一种基于SMT的智能体验证工具,可将SMT支持的证明失败本地化、可检查且可修复。Schwarz将失败的验证转化为约束本地化修复任务:程序点快照暴露边界处的已检查事实,局部引理让智能体提出缺失的证明步骤,感知理论的求解器策略引导智能体生成对求解器友好的公式,适用于数值、量化、内存和浮点约束。我们为C和Rust/Verus实现Schwarz,并在1475个任务上评估:在近期智能体验证工具的475个基准中,Schwarz解决95.2%的任务;在SV-COMP 2026 ReachSafety赛道的1000个任务(平均1427行代码)中,Schwarz解决91.5%的任务,而CPAchecker仅解决60.1%;消融实验与纯智能体基线的对比表明,感知求解器的修复方法有效且可扩展。

英文摘要

Agentic verification systems can often generate source-level specifications that look plausible, but plausibility is not enough: the verifier must still turn those specifications into SMT obligations that the solver can prove. When this step fails, current LLM-driven loops usually expose only a coarse verifier error, timeout, or unknown solver result. The model cannot tell whether the specification is wrong, a helper lemma is missing, the proof context contains irrelevant facts, or the obligation needs a different theory view. This paper presents Schwarz, an agentic verification harness that makes SMT-backed proof failure local, checkable, and repairable. Schwarz turns failed verification into obligation-local repair tasks: program-point snapshots expose checked facts at a boundary, local lemmas let the agent propose missing proof steps, and theory-aware solver policies guide the agent toward solver-friendly formulations for numeric, quantified, memory, and floating-point obligations. We implement Schwarz for C and Rust/Verus and evaluate it on 1,475 tasks. On 475 benchmarks from recent agentic verification tools, Schwarz solves 95.2% of the tasks. On 1,000 tasks from the SV-COMP 2026 ReachSafety track, averaging 1,427 LOC, Schwarz solves 91.5% of the tasks, compared with 60.1% for CPAchecker. Ablations and comparison with a pure-agent baseline show that solver-aware repair is effective and scalable.

Comments12 pages, 4 figures, 4 tables; preprint prepared in IEEE conference format

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑