AgentDV:用于硬件设计验证的闭环智能体AI
AgentDV: Closed-Loop Agentic AI for Hardware Design Verification
浏览论文内容
中文总结 AI 辅助
AgentDV是用于自动RTL验证的闭环智能体AI框架,通过三项关键技术优化,在三类LLM上显著提升了测试平台的通过率与代码、分支覆盖度。
中文摘要 AI 辅助
寄存器传输级(RTL)验证占据了现代片上系统(SoC)开发工作的很大一部分。然而,近期基于大语言模型(LLM)的验证代码生成往往无法产出可运行、与设计一致且能达到覆盖要求的测试平台。我们提出AgentDV,一个用于自动RTL验证环境生成的闭环智能体AI框架。AgentDV通过结合LLM引导的分析、测试平台构建、仿真、覆盖度测量和迭代优化,将单次LLM测试平台生成转变为基于工具的验证流程。该框架引入三个核心思路:1)可运行性过滤,用于拒绝无效的生成环境;2)基于控制/状态寄存器(CSR)的检查,以减少幻觉信号和不正确的预期行为;3)覆盖度引导的迭代,基于测得的验证缺口重新生成测试。我们在挑战被测设计(DUT)以及公开的OpenTitan外设和安全IP模块上,使用三种LLM对AgentDV进行评估。分析发现,直接单次提示在基准测试中无法产出有效的、能达到覆盖要求的环境。使用Claude Sonnet 4.6时,AgentDV在四个DUT上达到100%通过率,在所有DUT上平均通过率为80.9%;Llama和Qwen模型的平均通过率分别为58.7%和60.6%。此外,对于所考虑的基准测试,Claude Sonnet 4.6的平均行覆盖度为74.5%、分支覆盖度为88.4%,Llama的平均行覆盖度为69.1%、分支覆盖度为82.3%,Qwen的平均行覆盖度为64.9%、分支覆盖度为76.7%。
英文摘要
Register-transfer level (RTL) verification consumes a major part of modern system-on-chip (SoC) development effort. Yet, recent LLM-based verification-code generation often fails to produce runnable, design-consistent, and coverage-producing testbenches. We present AgentDV, a closed-loop agentic AI framework for automated RTL verification environment generation. AgentDV transforms single-shot LLM testbench generation into a tool-grounded verification pipeline by combining LLM-guided analysis, testbench construction, simulation, coverage measurement, and iterative refinement. The framework introduces three key ideas: 1) runnability filtering to reject invalid generated environments, 2) CSR-grounded checking to reduce hallucinated signals and incorrect expected behavior, and 3) coverage-guided iteration to regenerate tests based on measured verification gaps. We evaluate AgentDV using three LLMs on challenge DUTs and public OpenTitan peripheral and security IP blocks. From our analysis, we observed that direct single-shot prompting fails to produce a valid coverage-producing environment on benchmarks. AgentDV achieves 100% pass rate on four DUTs and an average of 80.9% pass rate on all DUTs using Claude Sonnet 4.6. Similarly, an average of 58.7% and 60.6% pass rate is achieved for Llama and Qwen models, respectively. In addition, an average of 74.5%, 69.1%, and 64.9% of line coverage and 88.4%, 82.3%, and 76.7% of branch coverage for the benchmarks under consideration for Claude Sonnet 4.6, Llama, and Qwen models, respectively.