SkillForge:结合验证循环的组合技能合成方法,用于生成形式化验证的 Dafny 程序
SkillForge: Compositional Skill Synthesis with Verification-in-the-Loop for Generating Formally Verified Dafny Programs
- Zhejiang University(浙江大学)
- Southeast University(东南大学)
- MIT(麻省理工学院)
- Weixin AI Lab(微信AI实验室)
- Renmin University(中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SKILLFORGE 是将形式化代码合成分解为原子技能库并结合验证循环协调的框架,在自然语言转 Dafny 程序任务上性能优于现有方法。
AI中文摘要:
从自然语言生成形式化验证的程序仍然具有挑战性:现有方法要么在单次传递中生成代码,验证失败时无补救措施,要么依赖开放式智能体推理,这种推理具有不确定性且不透明。我们提出 SKILLFORGE,这是一个将形式化代码合成分解为原子、可复用技能库的框架,每个技能针对特定子任务,如规范推断、主体合成、不变量生成、错误诊断或针对性修复,由提示模板、工具绑定和可判定成功准则定义。一个验证驱动的控制器协调这些技能:它将候选代码提交给 Dafny 验证器,将失败诊断为结构化类别,确定性地路由到适当的修复技能,并迭代直到形式化正确性得到证明或预算耗尽。在精心整理的自然语言到 Dafny 规范对基准上,SKILLFORGE 在 substantially 优于最先进的智能体方法(包括 ReAct 风格智能体、基于 MCTS 的修复和 RL 引导验证)以及传统迭代基线,同时需要更少的 token 和更低的延迟。消融研究证实每个技能都有可衡量的贡献,且控制器收敛迅速,大部分程序在首次尝试时即得到验证。
英文摘要:
Generating formally verified programs from natural language remains challenging: existing approaches either produce code in a single pass without recourse when verification fails, or rely on open-ended agentic reasoning that is non-deterministic and opaque. We introduce SKILLFORGE, a framework that decomposes formal code synthesis into a library of atomic, reusable skills, each targeting a specific subtask such as specification inference, body synthesis, invariant generation, error diagnosis, or targeted repair, and defined by a prompt template, tool binding, and decidable success criterion. A verification-driven harness orchestrates these skills: it submits candidates to the Dafny verifier, diagnoses failures into structured categories, deterministically routes to the appropriate repair skill, and iterates until formal correctness is proved or a budget is exhausted. On a curated benchmark of natural language to Dafny specification pairs, SKILLFORGE substantially outperforms both state-of-the-art agentic approaches (including ReAct-style agents, MCTS-based repair, and RL-guided verification) and traditional iterative baselines, while requiring fewer tokens and lower latency. Ablation studies confirm that every skill contributes measurably, and the harness converges rapidly with the majority of programs verified on the first attempt.