arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

未言明即不安全?基于大语言模型的RTL代码生成中的隐式安全义务

Unsaid, Unsafe? Implicit Security Obligations in LLM-Based RTL Code Generation

Guang Yang, Xing Hu, Xiang Chen, Xin Xia

arXiv 2608.26588首次发表:更新:

AI 中文总结

该研究针对LLM生成RTL代码的隐式安全问题,构建了SECRTL-GEN基准,提出神经符号框架RTL-Obliger,可显著提升LLM生成RTL的安全与功能通过率。

AI 中文摘要

大语言模型(LLM)生成寄存器传输级(RTL)代码的功能正确性正快速提升。然而,针对LLM生成代码的安全研究主要集中在软件领域,软件缺陷可在部署后修复,但不安全的RTL一旦流片进入硅片便无法补救。我们构建了SECRTL-GEN,这是一个基于真实片上系统知识产权(SoC IP)的多语言资源访问安全基准:包含5个常见弱点枚举(CWE)家族和4种硬件描述语言(HDL,即Verilog、SystemVerilog、VHDL、Python)的392个任务,每个任务均配有黑盒功能测试平台和安全测试平台。功能规范刻意省略了安全义务,与实际中安全义务常被排除在功能文档之外的情况一致。对5款前沿LLM的实证研究显示存在显著差距:在普通提示词下,它们通过功能测试的比例约为73%-79%,但通过安全测试的比例仅为14%-35%,且功能更强的模型并不更安全。加入CWE知识可提升安全性,而无辅助的自我思考帮助有限,两种面向安全的提示词均会降低功能通过率,表明瓶颈在于规范中缺失的弱点意识,而非编写防御性RTL的能力不足。我们提出RTL-Obliger,一个神经符号框架,用于推断这些隐式义务:LLM从规范中提取功能语义图,符号引擎随后将其与CWE模式本体匹配,以发现缓解措施证据缺口和信号级义务,最后LLM在这些义务约束下以功能保留的两阶段生成方式修改RTL。在5款模型和4种语言上,RTL-Obliger将平均全通过率从安全生成基线(SecV/RESCUE)的49.6%-51.4%提升至61.6%,且安全性和功能通过率均高于这些基线。

英文摘要

Large Language Models (LLMs) generate register-transfer-level (RTL) code with rapidly improving functional correctness. Security of LLM-generated code, however, has been studied mainly for software, where flaws can still be patched after deployment. Insecure RTL offers no such remedy once taped out into silicon. We construct SECRTL-GEN, a multi-language resource-access security benchmark grounded in real SoC IP: 392 tasks over five CWE families and four HDLs (Verilog, SystemVerilog, VHDL, and Python), each with black-box functional and security testbenches. Functional specifications intentionally omit security obligations, matching how obligations are often kept out of functional docs in practice. An empirical study of five frontier LLMs shows a sharp gap: under vanilla prompts they pass functional tests in about 73-79% of cases but security tests in only 14-35%, and stronger functional models are not safer. Adding CWE knowledge raises security, while unaided self-thinking helps less and both security-oriented prompts cut functional pass rates, showing that the bottleneck is missing weakness awareness in the specification, not an inability to write defensive RTL. We present RTL-Obliger, a neuro-symbolic framework that infers these implicit obligations. An LLM extracts a functional-semantic graph from the specification; a symbolic engine then matches it against a CWE pattern ontology to surface mitigation-evidence gaps and signal-level obligations; the LLM finally revises RTL under those obligations in a functionality-preserving two-stage generation. Across five models and four languages, RTL-Obliger raises mean all-pass from 49.6-51.4% (SecV/RESCUE) to 61.6%, with higher security and functional rates than these secure-generation baselines.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑