arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12635cs.AR

GateTruth:通过变异测试审计RTL设计基准的严谨性

GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing

Meet Bhadra

AI总结:

GateTruth是用于审计RTL基准测试平台严谨性的变异测试引擎,实验显示多数RTLLM v2.0基准测试平台未达95%变异体杀伤阈值,4096-token输出上限会影响模型排名,其成果推动将变异体杀伤认证设为RTL基准的标准要求。

AI中文摘要:

用于评估大型语言模型在寄存器传输级(RTL)硬件设计上表现的基准已迅速增多,但尚无任何基准报告采用变异测试这一成熟的硬件验证技术来量化测试平台质量,以检验自身的测试平台是否可信。从未失效的测试平台并不能证明设计正确,它可能只是从未触发实际存在缺陷的逻辑。我们提出GateTruth,一种用于审计RTL基准测试平台严谨性的变异测试引擎与方法:向参考设计注入一组确定性、带种子的语义变异体,并测量测试平台捕获的变异体比例。我们针对我们自己的68任务双轨参考套件验证了该方法——包含60项从规范到RTL生成的任务和8项智能体修复任务,通过固定的确定性综合到时序流程评分,且正确性被强制为严格的“门”,验证了60项A轨测试平台中有46项在可重复的顺序执行下能杀死至少95%的注入变异体;我们披露了其余14项未达标的原因,包括测试平台因修改以通过该“门”而产生的古德哈特定律效应。随后我们将同一未修改的引擎应用于广泛采用的外部基准RTLLM v2.0:在46项可审计设计中,72%低于我们自身套件所要求的95%阈值,且3项得分直接为0%。对NVIDIA的CVDP基准进行类似审计在结构上不可行:其公开版本未提供参考解决方案,缺失了变异测试所需的黄金RTL。审计我们自身的工具还发现了第二个结果:初始统一的4096-token输出上限静默截断了7个被评估模型中的3个,而以16384-token重新运行时,其中一个模型从第五名升至第一名。我们认为变异体杀伤认证应成为通用RTL生成基准的标准报告要求。

英文摘要:

Benchmarks for evaluating large language models on register-transfer-level (RTL) hardware design have proliferated rapidly, yet none reports having applied mutation testing, an established hardware-verification technique for quantifying testbench quality, to ask whether its own testbenches are trustworthy. A testbench that never fails is not evidence of a correct design; it may simply never stimulate the logic that is actually broken. We introduce GateTruth, a mutation-testing engine and methodology for auditing RTL benchmark testbench rigor: inject a deterministic, seeded set of semantic mutants into a reference design and measure what fraction the testbench catches. We validate the methodology against our own 68-task, dual-track reference suite -- 60 specification-to-RTL generation tasks and 8 agentic-repair tasks, scored through a pinned, deterministic synthesis-to-timing flow with correctness enforced as a strict gate -- certifying that 46 of 60 Track A testbenches kill at least 95% of injected mutants under sequential, reproducible execution; we disclose why the other 14 do not, including a Goodhart effect on testbenches revised to pass this gate. We then point the same engine, unmodified, at RTLLM v2.0, a widely adopted external benchmark: of 46 auditable designs, 72% fall below the 95% floor our own suite is held to, and three score 0% outright. A comparable audit of NVIDIA's CVDP benchmark is structurally impossible: its public release withholds reference solutions, removing the golden RTL mutation testing requires. Auditing our own instrument also surfaced a second finding: an initially uniform 4096-token output cap silently truncated three of seven evaluated models, and re-running at 16,384 tokens moved one model from fifth place to first. We argue mutation-kill certification should become a standard reporting requirement for RTL-generation benchmarks generally.

↑