发表机构
Fortinet, Inc.; Sony Corporate of America; Google LLC(飞塔公司; 美国索尼公司; 谷歌有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI生成的RTL验证计划在提供方模式与执行边界间失败的问题,提出可审计框架SecTB-RTL,通过自编安全回归代理验证,揭示模式接受不等于执行有效,并发布基准与治理控制。
AI 中文摘要
AI生成的RTL验证计划可能满足提供方模式,却在可信执行的边界处失败。我们提出SecTB-RTL,一个可审计的框架,涵盖31个任务和124个自编硬件安全回归测试。一个确定性的非AI基线在递增的资源限制下分别杀死了36、75和78个变异体。第一次确认性运行(C1-R2)在模型执行前失败,因为提供方拒绝了其响应模式。在未查看结果的情况下仅进行模式修复后,一个单独冻结的后续运行(C1-R3)完成了1,860次调用。提供方接受了1,857个响应,但只有九个通过了生产语义验证器。生成规则和执行规则不匹配。因此,我们将该运行保留为工具验证事件,并报告无提示效应估计。该事件表明,提供方或模式的接受并不确立执行有效性。编译和覆盖率仅是诊断手段;确切保存的工件必须通过完整生产路径。随后的后续运行被排除,因为它未满足预注册的证据完整性门控,仅被视为未来工作。我们发布了基准、故障保留契约、事件溯源以及防止基础设施行为被误报为模型行为所需的治理控制。
英文摘要
AI-generated RTL verification plans can satisfy a provider schema yet fail at the boundary to trusted execution. We present SecTB-RTL, an auditable framework covering 31 tasks and 124 authored hardware-security regressions. A deterministic non-AI baseline killed 36, 75, and 78 mutants at increasing resource limits. The first confirmatory run (C1-R2) failed before model execution because the provider rejected its response schema. After a schema-only repair made without viewing outcomes, a separately frozen follow-up run (C1-R3) completed 1,860 calls. The provider accepted 1,857 responses, but only nine passed the production semantic validator. The generation and execution rules did not match. We therefore preserve the run as an instrument-validation incident and report no prompt-effect estimate. This incident shows that provider or schema acceptance does not establish execution validity. Compilation and coverage are only diagnostics; the exact saved artifact must pass the full production path. A subsequent follow-up is excluded because it did not satisfy the preregistered evidence-completeness gate and is treated only as future work. We release the benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being misreported as model behavior.
CommentsCyber-AI