arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27067cs.LGcs.CL

ChipMEM:面向EDA智能体的验证接地记忆

ChipMEM: Verification-Grounded Memory for EDA Agents

Abdulrahman AlRabah, Joshua Mabry, Dilek Hakkani-Tür, Abdussalam Alawini, Hamid Shojaei, Kartik Hegde, Sandesh Adhikary

首次发表
浏览论文内容

中文总结 AI 辅助

提出ChipMEM,一种面向EDA智能体的验证接地记忆层,结合跨任务程序性记忆与贝叶斯统计引导,仅在通过综合、仿真或形式化检查后存储技能,提升RTL优化与测试平台生成性能,并实现技能向未见任务的迁移。

中文摘要 AI 辅助

基于大语言模型(LLM)的智能体利用电子设计自动化(EDA)工具,在综合和验证反馈下生成并修改寄存器传输级(RTL)设计。近期方法通过从执行轨迹中提炼可复用技能或基于EDA工具导出的奖励进行训练,来从这些反馈中学习。这两种方法通常都在产生经验的同一任务上进行评估。在同一任务上反复访问基准反馈可能奖励任务特定的修改,而非创建可迁移的复用知识。我们提出ChipMEM,一种面向EDA智能体的验证接地记忆层。它结合了跨任务的程序性记忆与轨迹内的统计引导。其程序性组件仅在技能通过综合、仿真或形式化检查后才进行提炼和存储,而非依赖模型自我评估。贝叶斯组件维护对工具调用结果的分层Beta估计,并对在类似错误下成功的恢复策略进行排序。一个通用适配器将相同的记忆接口应用于RTL优化和测试平台生成智能体,同时保留各领域的工具和验收标准。我们在训练任务上衡量性能,并评估所学技能是否迁移到未见任务。在RTLRewriter-Bench上,在匹配的模型和工具设置下,ChipMEM在39/54个评分设计中产生等价通过输出,而无记忆时为35/54;在49个设计的短套件中,平均面积改进为8.69%,对比5.66%。在留出的CVDP任务上,使用冻结程序库的ChipMEM在每次设置的单次评估中实现20/20的接受结果,而无记忆时为18/20。

英文摘要

Large language model (LLM)-based agents use Electronic Design Automation (EDA) tools to generate and revise register-transfer-level (RTL) designs under synthesis and verification feedback. Recent methods learn from this feedback by distilling reusable skills from execution traces or by training on rewards derived from EDA-tools. Both methods are typically evaluated on the tasks that produced the experience. Repeated access to benchmark feedback on the same task can reward task-specific revision rather than creating reusable knowledge that transfers. We introduce ChipMEM, a verification-grounded memory layer for EDA agents. It combines cross-task procedural memory with within-trajectory statistical guidance. Its procedural component distills and stores a skill only after it passes synthesis, simulation, or formal checks, rather than relying on model self-assessments. A Bayesian component maintains hierarchical Beta estimates over tool-call outcomes and ranks recovery strategies that succeeded under comparable errors. A common adapter applies the same memory interface to RTL optimization and testbench-generation agents while preserving each domain's tools and acceptance criteria. We measure performance on training tasks and evaluate whether learned skills transfer to unseen tasks. On RTLRewriter-Bench, under matched model and tool settings, ChipMEM produces equivalence-passing outputs on 39/54 scored designs versus 35/54 without memory; on the 49-design short suite, mean area improvement is 8.69% versus 5.66%. On held-out CVDP tasks, ChipMEM with a frozen procedural library achieves 20/20 accepted outcomes versus 18/20 without memory in a single evaluation per setting.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • NVIDIA(英伟达)
  • Cadence(楷登电子)

机构由 AI 辅助整理,请以论文原文为准。

↑