arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

芯片设计验证中强化学习的验证奖励模型

Verification Reward Model for Reinforcement Learning in Chip Design Verification

Shashank Chaurasia

arXiv 2609.22347首次发表:更新:

发表机构

MooresLab AI(MooresLab AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出验证奖励模型(VRM)框架,用于训练语言模型创建和修复芯片验证工件,通过版本化契约和确定性奖励合成提升验证质量,并规划了评估实验。

AI 中文摘要

我们提出了一种验证奖励模型(VRM)框架,用于训练语言模型以创建和修复芯片验证工件。核心对象是一个版本化的验证契约,该契约将需求、允许的激励、观测边界、参考行为、评估预算和验收标准绑定在一起。编译器、模拟器、形式化验证、变异测试、覆盖率和专家评审的证据被转换为可审计的记录。确定性验收检查保持在所学模型之外。一个学习到的结果模型根据规格说明、生成的工件以及明确掩蔽的部分证据来预测昂贵的未来证据;一个语义批评者识别出有证据支持的弱点;一个确定性奖励合成器将所得质量向量转换为任务条件下的训练奖励。扩展功能包括成对的干净/故障干预、边际故障发现奖励、不确定性感知的评估调度以及隔离的跨域学习循环。评估强调独立验证的故障检测、误报、跨设计族的泛化能力以及达到指定质量水平的总成本。我们描述了一个在小型、开放工具兼容基准上使用轻量级模型的首次实验,并通过特定功能的资格认证实现了完整的UVM能力。这是一篇立场论文:我们规定了框架和将对其进行测试的实验,并且我们不报告任何训练、EDA或硅片结果。

英文摘要

We propose a Verification Reward Model (VRM) framework for training language models to create and repair chip verification artifacts. The central object is a versioned verification contract that binds requirements, permissible stimulus, observation boundaries, reference behavior, evaluation budgets, and acceptance criteria. Compiler, simulator, formal, mutation, coverage, and expert-review evidence are converted into auditable records. Deterministic acceptance checks remain outside the learned model. A learned outcome model predicts expensive future evidence from the specification, the generated artifact, and explicitly masked partial evidence; a semantic critic identifies evidence-supported weaknesses; a deterministic reward composer translates the resulting quality vector into task-conditioned training rewards. Extensions include paired clean/fault interventions, marginal fault-discovery rewards, uncertainty-aware evaluation scheduling, and a quarantined cross-domain learning loop. Evaluation emphasizes independently validated fault detection, false alarms, generalization across design families, and total cost to reach a specified quality level. We describe a first experiment on a small, open-tool-compatible benchmark with lightweight models, with full UVM capability admitted through feature-specific qualification. This is a position paper: we specify the framework and the experiments that would test it, and we report no training, EDA, or silicon results.

CommentsPosition paper; 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑