HE-Guardrail:面向加密大语言模型推理的同态防护栏,抵御越狱攻击
HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference
- Pohang University of Science and Technology (POSTECH)(浦项科技大学)
- Inha University(仁荷大学)
- Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对HE-LLM推理中恶意客户端越狱攻击难以检测的问题,提出HE-Guardrail框架,在加密数据上同态运行防护栏并控制响应返回,实验复现明文防护栏决策,实现安全-效率-效用权衡。
AI中文摘要:
同态加密(HE)已成为隐私保护机器学习(PPML)的一种有前景的方法,能够直接对加密数据进行计算。在基于HE的PPML中,客户端向服务器提交加密输入,服务器在不访问底层明文的情况下评估大语言模型(LLM)等模型。然而,我们发现了该设置中的一个关键安全漏洞:HE-LLM推理容易受到提交对抗性提示(如越狱攻击)的恶意客户端的攻击。保护良性客户端的同一机密性也阻止了服务器检查传入的提示或生成的响应,使得对抗性尝试难以被检测或阻止,并可能使成功的攻击对服务器完全不可见。为了解决这一漏洞,我们提出了HE-Guardrail,一个完全在加密数据上评估防护栏机制并同态控制目标模型响应是否返回给客户端的框架。我们用三种具有代表性的防护栏——Llama Guard、JBShield和GradSafe——实例化了HE-Guardrail。我们的结果表明,HE-Guardrail在加密域中紧密复现了相应明文防护栏的决策,具有不同的安全-效率-效用权衡。
英文摘要:
Homomorphic encryption (HE) has emerged as a promising approach to privacy-preserving machine learning (PPML), enabling computation directly over encrypted data. In HE-based PPML, a client submits an encrypted input to the server, which evaluates models such as large language models (LLMs) without access to the underlying plaintext. However, we identify a critical security vulnerability in this setting: HE-LLM inference is vulnerable to malicious clients that submit adversarial prompts, such as jailbreak attacks. The same confidentiality that protects benign clients also prevents the server from inspecting incoming prompts or generated responses, making adversarial attempts difficult to detect or block and potentially allowing successful attacks to remain entirely invisible to the server. To address this vulnerability, we propose HE-Guardrail, a framework that evaluates guardrail mechanisms entirely over encrypted data and homomorphically controls whether the target-model response is returned to the client. We instantiate HE-Guardrail with three representative guardrails - Llama Guard, JBShield, and GradSafe. Our results show that HE-Guardrail closely reproduces the decisions of the corresponding plaintext guardrails in the encrypted domain, with distinct security-efficiency-utility trade-offs.