逃逸对齐:Best-of-N 越狱的物理陷阱模型
Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking
AI总结:
本文提出物理陷阱模型,用四个可解释参数刻画Best-of-N越狱的完整攻击面,实现跨模型和温度的外推预测。
AI中文摘要:
Best-of-$N$ 越狱(BoN)通过抽取不安全提示的 $N$ 个独立增强版本,并对每个版本采样 $M$ 个补全,从而绕过对齐模型的防护机制。以往研究表明,攻击成功率(ASR)似乎随 $N$ 呈幂律分布,但我们对此提出质疑。指数随 $N$ 漂移,存在指数交叉,这是对抗性数据集的有限尺寸伪影。目前鲜有工作探索整个双预算($N, M$)攻击面及其对生成温度 $T$ 的依赖性。我们引入一个简单的势垒模型,其中每个提示具有基线安全水平,每个增强具有随机热激活势垒。然后,由四个数字(每个都基于可解释的安全机制)决定整个($N, M$)攻击面。这些数字可将预测从 $N \leq 100$ 外推至 $N = 10^4$,将五个不同模型折叠到同一标度函数上,并可从拟合温度预测不同温度下的 ASR。
英文摘要:
Best-of-$N$ jailbreaking (BoN) bypasses safeguards of aligned models by drawing $N$ independent augmentations of an unsafe prompt and sampling $M$ completions of each. Previous works have shown that the attack success rate (ASR) seems to follow a power-law in $N$, which we challenge. The exponent drifts with $N$, with an exponential crossover which is a finite-size artifact of the adversarial dataset. Little work has been done to explore the entire two-budget ($N, M$) attack surface as well as its dependence on the generation temperature $T$. We introduce a simple barrier model where each prompt has a baseline safety level and each augmentation a random thermally activated barrier. Then four numbers, each backed by an interpretable safety mechanism, determine the entire ($N, M$) attack surface. They extrapolate predictions from $N \leq 100$ to $N = 10^4$, collapse five distinct models on the same scaling function and predict ASR at different temperatures from the one they were fitted at.