arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12762cs.AI

PROVE-RT:使用大语言模型为实时系统生成机械化定理证明器脚本

PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs

  • Florida International University(佛罗里达国际大学)
  • University of South Florida(南佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

Sadat Shahriyar, Shareef Ahmed, Abdullah Al Arafat

AI总结:

PROVE-RT是一种LLM辅助框架,通过依赖感知草图、PROSA文档检索等步骤生成PROSA/ROCQ脚本,在1191篇实时系统论文构建的语料库上,其机械化可调度性分析的成功率达44.7%,优于直接提示的LLM。

AI中文摘要:

可调度性分析是验证实时系统的关键,但现有测试多通过纸笔证明开发,难以扩展、验证和维护。PROSA/ROCQ中的机械化验证提供了严格替代方案,但手动构建此类证明需要大量领域专业知识和证明工程工作量。大语言模型(LLM)在各类任务中取得的成功使其成为为机械化定理证明器生成PROSA/ROCQ脚本的有潜力候选。然而,最先进的LLM往往缺乏正确使用PROSA建模抽象和证明模式所需的PROSA特定知识。本文提出PROVE-RT,一种LLM辅助框架,用于为实时系统文献中的可调度性分析生成PROSA/ROCQ脚本,该框架通过依赖感知的非正式草图、从处理后的PROSA文档检索、分阶段骨架生成和证明完成来指导生成。我们从1191篇实时系统论文构建了面向机械化的语料库,包含13134个带依赖信息的非正式草图。在精心整理的评估集上,最先进的LLM的直接提示无法可靠生成有效的PROSA机械化内容,而PROVE-RT的成功率达到44.7%。这些结果表明,检索引导和分阶段的LLM辅助可改善PROSA/ROCQ中可调度性分析的自动化机械化。

英文摘要:

Schedulability analysis is essential for certifying real-time systems, but existing tests are often developed through pen-and-paper proofs that are difficult to scale, validate, and maintain. Mechanized verification in PROSA/ROCQ offers a rigorous alternative, yet manually constructing such proofs requires substantial domain expertise and proof-engineering effort. Recent successes of large language models (LLMs) across a wide range of tasks make them promising candidates for generating PROSA/ROCQ scripts for mechanized theorem provers. However, state-of-the-art LLMs often lack the PROSA-specific knowledge required to correctly use its modeling abstractions and proof patterns. This paper introduces PROVE-RT, an LLM-assisted framework for generating PROSA/ROCQ scripts to mechanize schedulability analyses in real-time systems literature. PROVE-RT guides generation through dependency-aware informal sketches, retrieval from processed PROSA documentation, staged skeleton generation, and proof completion. We construct a mechanization-oriented corpus from 1, 191 real-time systems papers, containing 13, 134 informal sketches with dependency information. On a curated evaluation set, direct prompting of state-of-the-art LLMs fails to reliably generate valid PROSA mechanizations, whereas PROVE-RT achieves a success rate of 44.7%. These results show that retrieval-guided and staged LLM assistance can improve automated mechanization of schedulability analysis in PROSA/ROCQ.

↑