arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25335cs.ARcs.PL

GRADE-RTL:超越编译的LLM生成RTL评估

GRADE-RTL: Evaluating LLM-Generated RTL Beyond Compilation

Hepziba Susan, Shivaranjani G. R., Malik Imran, Muhammad Rashid, Sumathi Gokulanathan, Zain Ul Abideen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出GRADE-RTL框架,通过五项检查评估LLM生成的RTL代码,超越编译验证,区分结构、功能与实现质量,实验显示端到端成功率0%至70%,并揭示功能等价RTL在物理实现上的显著差异。

中文摘要 AI 辅助

大语言模型(LLM)能够根据自然语言规范生成寄存器传输级(RTL)代码,但仅靠编译无法确定结构完整性、功能正确性或实现效率。本文提出了一个超越编译的LLM生成RTL评估框架,命名为GRADE-RTL。通过使用固定提示和验证设置,GRADE-RTL应用五项检查,包括端口签名、编译、细化、模块完整性和与可信参考RTL的功能等价性。借助紧凑的指标,我们识别候选代码失败的位置,并区分端到端成功与下游综合的资格。我们在十种边缘相关知识产权设计上评估了九个通用型和RTL专用型LLM,预算为三次生成尝试并带有失败导向的反馈。端到端成功率在评估模型中从0%到70%不等,失败超出编译范围,涉及层次结构解析、逻辑不完整和行为不匹配。FPGA实现和65纳米ASIC综合结果进一步表明,功能等价的RTL在资源使用、时序、面积和功耗方面可能存在显著差异。一个PID控制器布局布线案例研究说明了不同RTL实现的物理设计后果。GRADE-RTL通过区分结构有效性、行为正确性和实现质量,为比较LLM生成的硬件描述提供了实用基础。

英文摘要

Large language models (LLMs) can generate register-transfer-level (RTL) code from natural-language specifications, but compilation alone does not establish structural completeness, functional correctness, or implementation efficiency. This paper presents a framework for evaluating LLM-generated RTL beyond compilation, which we named GRADE-RTL. Using fixed prompts and validation settings, GRADE-RTL applies five checks, such as Port Signature, Compilation, Elaboration, Module Completeness, and Functional Equivalence against trusted reference RTL. With the help of compact metrics, we identify where candidates fail and distinguish end-to-end success from eligibility for downstream synthesis. We evaluate nine general-purpose and RTL-specialized LLMs on ten edge-relevant intellectual property designs under a budget of three generation attempts with failure-directed feedback. End-to-end success ranges from 0% to 70% across the evaluated models, with failures extending beyond compilation to hierarchy resolution, incomplete logic, and behavioral mismatch. FPGA implementation and 65 nm ASIC synthesis results further show that functionally equivalent RTL can differ substantially in resource use, timing, area, and power. A PID-controller place-and-route case study illustrates the physical-design consequences of different RTL implementations. GRADE-RTL provides a practical basis for comparing LLM-generated hardware descriptions by separating structural validity, behavioral correctness, and implementation quality.

发表机构

  • Vellore Institute of Technology(维洛尔理工学院)
  • Queen’s University Belfast(贝尔法斯特女王大学)
  • Umm Al-Qura University(乌姆古拉大学)
  • University of Idaho(爱达荷大学)

机构由 AI 辅助整理,请以论文原文为准。

↑