arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19189cs.ARcs.CL

CovR:通过推理引导的强化学习实现覆盖率感知的硬件验证

CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning

Manar Abdelatty, Maryam Nouh, Sherief Reda

首次发表
浏览论文内容

中文总结 AI 辅助

CovR提出一种结合自我反思与仿真反馈的智能体框架,通过强化学习优化覆盖率驱动的测试平台生成,在多个基准上显著提升覆盖率并发现更多隐藏故障。

中文摘要 AI 辅助

设计验证仍然是硬件开发中资源最密集的阶段之一,通常消耗总设计工作量的高达70%。尽管近期已有工作探索使用大型语言模型(LLMs)来自动化测试平台生成,但大多数现有方法仅狭隘地关注功能正确性,忽视了覆盖率质量这一关键方面。为弥补这一差距,我们提出了CovR,一个用于自动化测试平台生成的智能体框架,它将自我反思循环与基于仿真的反馈相结合,以最大化覆盖率。利用这一流程,我们借助一个强大的教师模型,构建了一个包含16,514个自然语言规范RTL推理测试平台元组的大规模数据集,从而实现了覆盖率感知的监督。在此基础上,我们提出了一种专为覆盖率驱动的测试平台生成而设计的强化学习(RL)框架,利用从仿真和覆盖率反馈中获得的工具衍生奖励来优化学生模型。实验结果表明,CovR微调模型在VerilogEval和RTLLM V2.0上达到了93.81%的cov@10,在CVDP上达到了87.76%的cov@10,分别比最先进的方法高出7.97%和3.59%。此外,将微调模型重新部署回智能体细化流程中,进一步将VerilogEval和RTLLM V2.0上的cov@10提升至94.27%,CVDP上的cov@10提升至91.39%。而且,当作为插件式激励引擎集成到完整验证工作流中时,CovR将覆盖率提高了18.95%,突变检测分数提高了1.19%,同时揭示了4.46%的未检测故障,凸显了在基于LLM的硬件验证中优化覆盖率的重要性。

英文摘要

Design verification remains one of the most resource-intensive stages of hardware development, often consuming up to 70% of the total design effort. While recent work has explored using Large Language Models (LLMs) to automate testbench generation, most existing approaches focus narrowly on functional correctness, overlooking the critical aspect of coverage quality. To bridge this gap, we present CovR, an agentic framework for automated testbench generation that combines self-reflection loops with simulation-based feedback to maximize coverage. Using this pipeline, we construct a large-scale dataset of 16,514 natural specification RTL reasoning testbench tuples with a strong teacher model, enabling coverage-aware supervision. Building on this, we propose a reinforcement learning (RL) framework tailored for coverage-driven testbench generation, leveraging tool-derived rewards from simulation and coverage feedback to optimize a student model. Experimental results show that the CovR finetuned model achieves 93.81% cov@10 on VerilogEval and RTLLM V2.0, and 87.76% cov@10 on CVDP, outperforming state-of-the-art approaches by 7.97% and 3.59%, respectively. Furthermore, deploying the finetuned model back into the agentic refinement pipeline further improves cov@10 to 94.27% on VerilogEval and RTLLM V2.0 and 91.39% on CVDP. Moreover, when integrated as a plug-in stimulus engine for full verification workflows, CovR improves coverage by 18.95% and mutation detection score by 1.19%, while revealing 4.46% undetected failures, highlighting the importance of optimizing for coverage in LLM-based hardware verification.

发表机构

  • Brown University(布朗大学)

机构由 AI 辅助整理,请以论文原文为准。

↑