AI 中文总结
CircuitProver是基于Lean 4的智能体硬件验证框架,可自动转换硬件设计为Lean模型,通过积累可复用证明知识完成63项基准证明,比普通智能体更高效。
AI 中文摘要
现代集成电路(IC)日益复杂,使得功能验证成为主要瓶颈。主流的硬件形式化验证方法——模型检验,会分别验证每个设计实例,仅输出通过/失败结果,导致证明背后的推理过程被封装在求解器启发式算法中,且在相关设计间被重复构建。交互式定理证明虽能生成显式、可复用的证明产物,但将其应用于硬件领域仍以人工为主,需专家投入大量精力完成形式化、不变量发现及证明开发。本文提出CircuitProver,这是一个基于Lean 4的智能体验证框架,支持证明积累与参数化验证。CircuitProver可自动将参数化硬件设计及其自然语言规范转换为可执行的Lean 4模型,随后通过Lean的反馈迭代构建经机器检验的证明,以确认硬件代码符合规范。其证明轨迹与已验证定理会被提炼为可复用库,其中证明策略可指导后续智能体推理,已验证引理则支持在相关硬件验证任务间复用形式化证明。我们还推出首个用于评估智能体硬件定理证明的基准套件,涵盖多样化的参数化硬件设计、规范、证明任务及评估指标。在63项任务中,CircuitProver成功证明所有基准,而普通智能体解决了92.1%的任务,平均所需证明轮次为CircuitProver的两倍。消融研究表明,积累的证明知识可减少相关验证任务间的冗余证明构建,使证明长度缩短16.3%,验证时间减少23.2%。
英文摘要
Modern integrated circuits (ICs) are becoming increasingly complex, making functional verification a major bottleneck. The dominant hardware formal verification methodology, model checking, verifies each design instance separately and exposes only pass/fail results, so the reasoning behind a proof stays locked inside solver heuristics and is repeatedly reconstructed across related designs. Interactive theorem proving instead yields explicit, reusable proof artifacts, but applying it to hardware remains largely manual, demanding expert effort for formalization, invariant discovery, and proof development. In this paper, we present CircuitProver, an agentic Lean 4-based verification framework supporting proof-accumulation and parameterized verification. CircuitProver automatically translates parameterized hardware designs and their natural language specifications into executable Lean 4 models. It then iteratively constructs machine-checked proofs through Lean feedback to establish that the hardware code complies with the specification. The proving traces and verified theorems are distilled into reusable libraries, where proving strategies guide future agent reasoning and verified lemmas support formal proof reuse across related hardware verification tasks. We further introduce the first benchmark suite for evaluating agentic hardware theorem proving, covering diverse parameterized hardware designs, specifications, proof tasks, and evaluation metrics. Across 63 tasks, CircuitProver successfully proves all benchmarks, while a vanilla agent solves 92.1% of them and requires twice as many proof rounds on average. Ablation studies show that accumulated proof knowledge reduces redundant proof construction across related verification tasks, reducing proof length by 16.3% and verification time by 23.2%.