发表机构
NCE Chandi; IIT (BHU); IIT Patna(NCE Chandi; 印度理工学院(瓦拉纳西); 巴特那印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对本地部署大语言模型的输入端越狱防御开展审计,追溯防御失效的假设根源,在6款14B至35B参数的开源权重模型上,用100条来自40余公开源的越狱提示完成13800条评估记录的测试。
AI 中文摘要
通过Ollama等推理引擎本地部署的大语言模型(LLMs),不具备API服务模型中存在的审核与滥用检测机制,因此LLMs的安全性依赖于所使用的防御机制,而其有效性取决于设计时所依据的假设。本文对本地部署模型在越狱攻击下的防御机制进行审计,部分防御提供形式化保证(SmoothLLM、Erase-and-Check、Sequential Monitors),其余则依赖经验检测结果(Semantic Smoothing、Self-Denoised Smoothing、Perplexity Filtering)。本文未仅观察到防御失效,而是将每一次失效追溯至具体假设:针对每种防御,提取其依赖的条件,推导违反该条件应产生的经验模式,并在6个开源权重模型(参数规模14B至35B)上测试该预测,所用语料来自40余个公开来源的100条越狱提示,共产生13800条评估记录。
英文摘要
Locally deployed Large Language Models (LLMs) via inference engines such as Ollama run without the moderation and abuse detection present in API-served models. Therefore, the safety of LLMs depends on the defense mechanisms used, and their effectiveness depends on the assumptions on which they were designed. This paper does an audit of defense mechanisms under jailbreak attacks on locally deployed models. Some defenses provide formal guarantees (SmoothLLM, Erase-and-Check, Sequential Monitors), while others rely on empirical detection results (Semantic Smoothing, Self-Denoised Smoothing, Perplexity Filtering). Instead of merely observing that defenses fail, we trace each failure back to the specific assumption: for every defense, we extract the condition it relies on, derive the empirical pattern a violation should produce, and test that prediction on six open-weight models (14B to 35B parameters) with a corpus of 100 jailbreak prompts taken from more than 40 public sources, totalling 13,800 evaluation records.