Reflex-Guard:一种使用稠密语义嵌入的低延迟大语言模型提示安全护栏
Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings
- Graduate School of Informatics, Osaka Metropolitan University(大阪公立大学情报学研究科)
- BRAC University(BRAC大学)
- Bangladesh University of Engineering and Technology (BUET)(孟加拉工程技术大学)
- New Mexico State University(新墨西哥州立大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
Reflex-Guard是一种本地运行的轻量型LLM提示安全护栏,采用稠密语义嵌入与快速分类器,延迟低至37.6毫秒,召回率达95.9%,性能优于现有基线,可高效检测各类越狱攻击。
中文摘要 AI 辅助
实际应用中的大语言模型(LLM)常面临精心设计的提示绕过安全控制的风险,现有护栏方法如LLM作为评判者、基于云的安全API虽能检测不安全内容,但通常会为每个请求增加约250-900毫秒的延迟,对于通常需要在100毫秒内响应的实时应用来说延迟过高;此外,将用户提示路由至外部审核端点会引发严重的数据隐私问题。本文介绍了Reflex-Guard,一种可本地运行的轻量型护栏,它采用感知越狱的预处理、紧凑的Sentence-Transformer嵌入以及7个快速二分类器,这些组件结合实现了高精度的提示安全过滤,且延迟远低于现有解决方案。通过对来自5个互补来源的30568个样本构成的策略性平衡数据集进行系统评估,研究表明Reflex-Guard对有害提示的召回率达95.9%,端到端延迟为37.6毫秒,其速度快于现有基线,包括延迟255毫秒的Llama Guard 2和延迟723毫秒的SafeDecoding;使用默认阈值时,它可检测100%的GCG后缀攻击和Base64编码的提示,而对于DrAttack结构化提示,需将阈值降至0.03以实现最优检测,因其会产生独特的概率分布。Reflex-Guard的反射效率得分(RES)最高达16.79,显著优于Llama Guard 2(11.90)和SafeDecoding(9.80),该分析提供了实用的部署建议,且显示不同攻击类型在嵌入概率空间中占据不同区域。
英文摘要
Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are able to detect unsafe content. However, they often add a delay of about 250-900 ms to each request. This delay is too high for real-time applications, when the system usually needs to respond in less than 100 ms. Furthermore, routing user prompts through external moderation endpoints raises significant data privacy concerns. This paper introduces Reflex-Guard, a lightweight guardrail that runs locally. It uses jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers. Together, these components enable high-accuracy prompt safety filtering with much lower latency than existing solutions. Through systematic evaluation on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, we demonstrate that Reflex-Guard achieves 95.9% recall on harmful prompts at 37.6 ms end-to-end latency. It is faster than existing baselines, including Llama Guard 2 at 255 ms and SafeDecoding at 723 ms. It can detect 100% of GCG suffix attacks and Base64-encoded prompts using the default threshold. However, DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection, as they produced a distinct probability distribution. Reflex-Guard achieves Reflex Efficiency Score (RES) scores up to 16.79, significantly outperforming Llama Guard 2 (11.90) and SafeDecoding (9.80). This analysis offers practical deployment advice and shows that different attack types occupy distinct regions in the embedding probability space.