arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过随机嵌入扰动越狱开放权重LLM

Jailbreaking Open-Weight LLMs via Random Embedding Perturbations

Abhinav Sudhakar Dubey, Scott Sirri, Vaggos Chatziafratis, C. Seshadhri

arXiv 2610.07125首次发表:更新:

发表机构

University of California, Santa Cruz(加州大学圣克鲁兹分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PEV攻击,通过向提示嵌入向量添加高斯噪声,以极低计算成本在六种开放权重LLM上实现快速越狱,揭示其安全漏洞。

AI 中文摘要

尽管开放权重模型在能力上取得了稳步进展,并在多个领域得到广泛采用,但其安全性仍然是一个重要关注点。其中一个关键特性是能够拒绝或回避有害、恶意或不敏感的提示。在本文中,我们揭示了六种常见不同规模的开放权重LLM在JailbreakBench基准数据集上存在安全漏洞,这些漏洞持续导致有害或不安全的响应。我们提出的攻击方法,即扰动嵌入向量(PEV),是一种简单且快速的“越狱”技术,其成本低于先前的方法,后者通常需要梯度计算、逐提示优化或修改模型的内部权重。PEV仅向提示的嵌入向量表示中添加独立的高斯噪声,无需进一步操作。为了生成不安全的响应,我们反复从该分布中采样加性噪声。在我们的实验中,我们观察到获得首次成功攻击的平均计算成本比之前的攻击低一个数量级。在每种测试模型上,新提示的首次成功越狱通常在一分钟内实现,并且PEV在JailbreakBench中的所有提示上对所有模型都生成了不安全的响应。没有其他测试方法能达到这样的结果,尽管它们运行时间更长。更广泛地说,我们认为理解LLM在嵌入向量扰动下的行为是一个重要的研究方向:虽然扰动构成了重大的安全风险,但它们也可以作为探索此类模型动态行为的有价值工具。

英文摘要

While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key feature is the ability to refuse or deflect harmful, malicious, or insensitive prompts. In this paper, we expose safety vulnerabilities across six common open-weight LLMs of various sizes that consistently lead to harmful or unsafe responses on the JailbreakBench benchmark dataset. Our proposed attack, Perturbed Embedding Vector (PEV), is a simple and fast "jailbreaking" technique that is cheaper than prior approaches, which typically require gradient computations, per-prompt optimizations, or altering internal weights of the models. PEV just adds independent Gaussian noise in the embedding vector representations of the prompt, with no need for further manipulations. To generate unsafe responses, we repeatedly sample additive noise from this distribution. In our experiments, we observe that the average compute cost to get the first successful attack is up to an order of magnitude less than previous attacks. The first successful jailbreak on a new prompt typically arrives within one minute on every tested model, and PEV generates unsafe responses across all models for all prompts in JailbreakBench. No other tested method achieves such results, despite them taking longer to run. More broadly, we believe that understanding the behavior of LLMs under perturbations in the embedding vectors is an important research direction: while perturbations constitute a major security risk, they can also serve as a valuable tool for exploring the dynamical behavior of such models.

Comments15 pages, 4 figures, Code: https://github.com/AbhinavDubey30/Perturbed-Embedding-Vectors

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑