arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33699cs.AIcs.SE

SpecRead:衡量语言模型是否理解硬件规格的基准

SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications

Feilian Huang

首次发表
浏览论文内容

中文总结 AI 辅助

SpecRead是一个衡量语言模型硬件规格理解能力的基准,通过385个问题分离理解与生成,并采用变异检测和一致性检查,初步评估显示模型在矛盾定位上表现较好但类别判定仍需提升。

中文摘要 AI 辅助

现有针对大型语言模型(LLM)在硬件设计领域的基准评估的是下游产物,如生成的RTL、断言或测试平台。当模型在此类基准上失败时,失败原因具有歧义性:它可能误读了规格,也可能理解了规格但未能编写出正确的代码。我们提出SpecRead,一个将规格理解能力与生成能力相分离的基准。SpecRead v2.1包含385个问题,覆盖10个开源OpenTitan IP模块:精确检索、跨章节推理、变异规格中的矛盾检测,以及规格-RTL一致性检查,另有82个对照组(41个干扰项,41个一致RTL)。类型4项目基于真实RTL变异构建;我们仅保留经Icarus Verilog仿真验证会改变可观察行为的变异。有规格与无规格的消融实验表明,这些问题需要依赖摘录而非仅凭训练记忆(在t1/t2子集上,无规格准确率为3/20),尽管对源文本的记忆可能仍有助于发现变异。作为使用小型模型的初步表征,Ministral-3B总体得分为33.2%(128/385;宏平均39.0%):检索任务55.2%,跨章节推理51.7%。在两项侧重矛盾的题型上,判定加定位的综合指标分别为48.0%(t3)和63.3%(t4),干扰项上的假阳性率为51.2%,一致RTL对照组上的假阳性率为100%。分层评分显示,模型定位矛盾的能力较好(定位准确率78.9%-81.6%),但在矛盾类别判定上得分较低(43.9%-49.7%)。一种结构化的“规则表”提示干预降低了除t2(持平)外所有题型上的准确率。SpecRead可通过确定性检查自动评分,在保守主评分下,灰色地带案例计为错误。该基准的类型3项目可通过变异注入重新生成,且完全基于公开来源构建。

英文摘要

Existing benchmarks for large language models (LLMs) in hardware design evaluate downstream artifacts such as generated RTL, assertions, or testbenches. When a model fails such a benchmark, the failure is ambiguous: it may have misread the specification, or it may have understood the specification and failed to write the code. We present SpecRead, a benchmark that isolates specification comprehension from generation ability. SpecRead v2.1 contains 385 questions over 10 open-source OpenTitan IP blocks: exact retrieval, cross-section reasoning, contradiction detection in mutated specifications, and spec-RTL consistency checking, plus 82 controls (41 distractor, 41 consistent-RTL). Type-4 items are built from real RTL mutations; we retain only mutations that Icarus Verilog simulation shows to change observable behavior. A with-spec vs. without-spec ablation suggests the questions require the excerpt, not training recall alone (without-spec accuracy 3/20 on the t1/t2 subset), though memorization of the source text may still help spot mutations. As an initial characterization with a small model, Ministral-3B scores 33.2% overall (128/385; macro average 39.0%): 55.2% on retrieval, 51.7% on cross-section reasoning. On the two contradiction-focused types, the verdict-plus-location measure gives 48.0% (t3) and 63.3% (t4), with a 51.2% false-positive rate on distractors and 100% on consistent-RTL controls. Layered scoring shows the model locates contradictions well (78.9-81.6% location accuracy) but scores lower on their category (43.9-49.7%). A structured "rule-table" prompting intervention lowers accuracy on every question type except t2 (tied). SpecRead is automatically scorable by deterministic checks, with gray-zone cases counted wrong under the conservative main scoring. The benchmark is regenerable for type-3 items via mutation injection, and built exclusively from public sources.

补充信息

↑