arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29870cs.SE

基于模型证明草图的正式模型构建指导

Formal Model Construction Guided by Model-Based Proof Sketches

  • National University of Singapore(新加坡国立大学)
  • IRIT - National Polytechnic Institute of Toulouse(图卢兹国立理工学院 IRIT)

机构由 AI 辅助整理,请以论文原文为准。

Hongshu Wang, Xinyue Zuo, Yufan Cai, Neeraj Kumar Singh, Yamine Ait Ameur, Jin Song Dong

AI总结:

针对现有自动形式化方法在修复时依赖验证反馈且可能顾此失彼的问题,提出基于模型证明草图的ProGS方法,通过树形结构引导LLM生成与修复,在27个系统基准上提升语法、验证与行为正确性。

AI中文摘要:

正式建模为系统正确性提供了强有力的保证,但开发和修复正式模型仍然劳动密集,且需要大量的逻辑和形式推理专业知识。近期基于LLM的自动形式化代理旨在通过生成候选正式模型并利用正式工具的反馈进行修订来减轻这一负担。然而,现有方法遵循生成-修复范式,其中修复由生成模型的验证失败驱动,因此高度依赖于反馈的粒度以及LLM的修复能力。结果,针对某一验证级别的修复可能会使另一级别的属性失效,这需要对完整的事件守卫集合进行推理。为解决这些限制,我们提出了基于模型证明草图的正式模型合成(ProGS),一种以基于模型的证明草图为中心的形式化方法。基于模型的证明草图将目标正式系统的证明结构表示为树。内部节点捕获情况分裂和归纳推理步骤,而叶节点对应于实现各个子目标的具体状态转换事件。ProGS利用LLM生成和修复这些草图,将验证失败映射回特定节点和子树,以提供结构化的迭代修复指导。我们在27个正式系统的基准上的评估显示,ProGS在语法有效性、演绎可验证性和行为正确性方面优于最先进的代理式正式建模方法,证明了围绕分层证明草图组织正式模型构建的益处。

英文摘要:

Formal modeling provides strong guarantees about system correctness, but developing and repairing formal models remains labor-intensive and requires substantial expertise in logic and formal reasoning. Recent LLM-based autoformalization agents seek to reduce this burden by generating candidate formal models and revising them using feedback from formal tools. However, the existing approaches follow a generate-and-repair paradigm, in which repairs are driven by verification failures of the generated model and therefore depend heavily on both the granularity of the feedback and the LLM's repair capability. As a consequence, a repair targeting one level of verification may invalidate properties at another level, which requires reasoning over the complete set of event guards. To address these limitations, we propose Proof-Sketch-Guided Formal Model Synthesis (ProGS), an autoformalization method centered on model-based proof sketches. A model-based proof sketch represents the proof structure of the target formal system as a tree. Internal nodes capture case splits and inductive reasoning steps, while leaf nodes correspond to concrete state-transition events that realize individual subgoals. ProGS uses LLMs to generate and repair these sketches, with verification failures mapped back to specific nodes and subtrees to provide structured guidance for iterative repair. Our evaluation on a benchmark of 27 formal systems shows that ProGS improves over state-of-the-art agentic formal modeling approaches in syntactic validity, deductive verifiability, and behavioral correctness, demonstrating the benefit of organizing formal model construction around hierarchical proof sketches.

↑