发表机构
Southwest University; Aarhus University; University of York; Centro de Informática, Universidade Federal de Pernambuco (UFPE); University of Lancaster(西南大学; 奥胡斯大学; 约克大学; 伯南布哥联邦大学信息中心; 兰卡斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出CAPRI工具,用于解决Isabelle证明修复中LLM修改未授权内容的问题,通过契约检查机制保障安全,在12个失败证明的实验中验证了不同工作流的修复效果。
AI 中文摘要
我们研究如何利用大语言模型(LLM)辅助发现Isabelle证明。Isabelle构建仅能确认提交的理论被接受,无法确保LLM仅修改了开发者授权的内容。我们提出CAPRI,一种感知契约的修复工作流,其中Isabelle负责检查证明,而独立检查器则执行机器可读的编辑契约。提示、提案、候选存储库、诊断、裁决和哈希值会被保留以供审计。我们针对四个开发项目中的12个失败证明评估了五种工作流,每个任务和条件下重复三次,共180次运行,产生138个有效修复。在Isabelle接受的144个终端候选中,有6个修改了受保护文本;这些均出现在可修改完整理论的迭代工作流中。仅证明体界面产生29/36的有效修复且无契约违规,而对应的完整理论工作流则为31/36。一次性修复产生22/36,后续前瞻性冻结的迭代工作流产生32/36;这些数字比较的是完整工作流而非单个机制。单独的事后OpenRouter活动在指定的Luna比较中未发现改进。带有匹配演示的Sol配置产生33/36的修复,而冻结的OpenAI Responses条件为29/36,但在单侧精确McNemar检验中差异无统计学意义(p=0.0625)。
英文摘要
We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theory is accepted, but not that an LLM changed only what the developer authorised. We present CAPRI, a contract-aware repair workflow in which Isabelle checks the proof and an independent checker enforces a machine-readable edit contract. Prompts, proposals, candidate repositories, diagnostics, verdicts, and hashes are retained for audit. We evaluate five workflows on twelve failed proofs from four developments, with three replicates per task and condition, giving 180 runs and 138 valid repairs. Of 144 terminal candidates accepted by Isabelle, six had modified protected text; all arose in iterative workflows that could edit a complete theory. A proof-body-only interface produced 29/36 valid repairs and no contract violations, compared with 31/36 for the corresponding full-theory workflow. One-shot repair produced 22/36, while a later prospectively frozen iterative workflow produced 32/36; these figures compare complete workflows rather than individual mechanisms. A separate post hoc OpenRouter campaign found no improvement in the designated Luna comparisons. A Sol configuration with matched demonstrations produced 33/36 repairs, compared with 29/36 in the frozen OpenAI Responses condition, but the difference was not statistically significant in a one-sided exact McNemar test ($p=0.0625$).
Comments17 pages, 1 figure, 7 tables. Submitted to SBMF 2026. Reproducibility artefact available on Zenodo