BEHAVE:功能行为建模使硬件设计与验证智能体能够自我改进
BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification
- Stanford University(斯坦福大学)
- Workato, Inc.(Workato公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
BEHAVE通过功能行为建模为硬件设计验证智能体提供可靠反馈,支持自我改进,将Qwen3.8-27B的RTL pass@1从55.0%提升至75.0%。
AI中文摘要:
开发用于硬件设计和验证的智能体需要可靠的正确性反馈。由于硬件规范可能允许具有不同延迟的正确实现,逐周期匹配设计和参考输出可能会拒绝有效的设计。为了解决这个问题,我们引入了BEHAVE,一个通过功能行为建模进行多轮联合硬件设计和验证的智能体框架。我们定义了行为IR(Behavior IR),将任务功能表达为可执行的行为模型,而不规定超出规范的实施时序。智能体迭代地开发寄存器传输级(RTL)设计和一个行为模型作为设计的验证参考。我们的评估器BEHAVE-Sim使用随机采样和求解器引导搜索生成的输入刺激,分别检查这两个工件与隐藏的黄金行为模型。因此,BEHAVE支持在任务允许的延迟和微架构下进行功耗、性能和面积(PPA)探索。在训练期间,相同的评估器从规范-行为对中提供可验证的强化学习(RL)奖励,而无需参考RTL。为了自我改进,智能体持续搜索与其能力差距相关的高级实现,构建并检查规范-行为对,并在扩展的任务池上进行训练。我们发布了BEHAVE-Train和BEHAVE-Eval,包含600个经过人工审查的规范-行为对,用于真实的硬件工作负载。从60个种子任务开始并获取100个新任务,自我改进将Qwen3.8-27B在BEHAVE-Eval上的RTL pass@1从55.0%提高到75.0%,达到了与使用540个任务池的RL相当的性能。
英文摘要:
Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid designs. To address this, we introduce BEHAVE, an agentic framework for multi-turn joint hardware design and verification through functional behavior modeling. We define Behavior IR to express task functionality as executable behavior models without prescribing implementation timing beyond the specification. The agent iteratively develops a register-transfer-level (RTL) design and a behavior model as the design's verification reference. Our evaluator, BEHAVE-Sim, checks both artifacts separately against a hidden golden behavior model using input stimuli generated by random sampling and solver-guided search. BEHAVE thus supports power, performance, and area (PPA) exploration across task-permitted latencies and microarchitectures. During training, the same evaluator provides verifiable reinforcement learning (RL) rewards from specification-behavior pairs without reference RTL. For self-improvement, the agent continually searches for high-level implementations relevant to its capability gaps, constructs and checks specification-behavior pairs, and trains on the expanded task pool. We release BEHAVE-Train and BEHAVE-Eval with 600 human-reviewed specification-behavior pairs for realistic hardware workloads. Starting from 60 seed tasks and acquiring 100 new tasks, self-improvement raises Qwen3.8-27B's RTL pass@1 on BEHAVE-Eval from 55.0% to 75.0%, reaching performance comparable to RL using a 540-task pool.