AI 中文总结
RTLGuard是一种轻量级师生防御方法,通过小规模干净教师模型、复合师生目标及特征对齐等,降低中毒RTL代码生成模型的攻击成功率,同时保留代码功能正确性与可综合性。
AI 中文摘要
大型语言模型(LLM)的快速发展正推动向自动寄存器传输级(RTL)代码生成的转变,使设计人员能够将高级规范转换为可综合的硬件。然而,对预训练(第三方)微调模型的依赖可能会引入关键的信任问题,因为这些模型的训练数据和适应过程通常是不透明的。因此,攻击者(甚至是模型提供者)可能会在微调过程中嵌入隐藏的后门威胁,使得在推理时,受害者用户给出看似良性的提示就能触发恶意行为,例如硬件特洛伊木马。在本文中,我们提出了RTLGuard,以缓解AI驱动的集成电路(IC)供应链中的此类信任问题。RTLGuard没有采用计算成本高昂的全参数微调,而是利用师生框架来净化中毒的RTL生成模型,具体包括:(1)在有限的可信RTL数据集上微调一个小规模的“干净”教师模型;(2)通过复合师生目标引导中毒的目标模型;(3)结合特征对齐和知识蒸馏来抑制恶意行为。我们在各种LLM架构上进行的实验表明,RTLGuard在显著降低攻击成功率(ASR)的同时,还能保持生成的RTL代码的功能正确性和可综合性。
英文摘要
The rapid advancement of large language models (LLMs) is driving a shift toward automated register transfer level (RTL) code generation, enabling designers to translate high-level specs. into synthesizable hardware. However, this reliance on pre-trained (3rd-party) fine-tuned models may introduce critical trust issues, as the training data and adaptation process of these models are often opaque. Thus, adversaries (even model providers) may embed hidden backdoor threats during fine-tuning, allowing malicious behavior, e.g., hardware Trojans, to be triggered by seemingly benign prompts given by victim user at inference time. In this paper, we introduce RTLGuard, to mitigate such a trust issue in AI-enabled IC supply chain. Rather than prohibitive computational cost of full-parameter retraining, RTLGuard leverages a teacher-student framework designed to sanitize compromised RTL generation models by (1) fine-tuning a small-scale, "clean" teacher model on a limited set of trusted RTL data, (2) guiding the poisoned target model via a composite teacher-student objective, and (3) incorporating feature alignment and knowledge distillation to suppress malicious behaviors. Our experiments across various LLM architectures demonstrate that RTLGuard significantly reduces the Attack Success Rate (ASR) while preserving the functional correctness and synthesizability of the generated RTL code.
Comments8 pages, 4 figures, 7 tables. Accepted at the IEEE/ACM International Conference on Computer-Aided Design (ICCAD 2026)