arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36956cs.CRcs.AI

黑盒大语言模型的受控解码攻击

Controlled Decoding Attacks on Black-Box LLMs

Jesson Wang, Shawn Li, Wei Yang, Franck Dernoncourt, Ryan A. Rossi, Charith Peris, Yue Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出受控解码攻击框架,通过文本接口的采样重建与选择性控制,实现黑盒大语言模型的高效越狱,并在多基准上取得最优效果。

中文摘要 AI 辅助

在生成过程中操纵下一个词元的概率可以绕过大型语言模型的安全对齐。然而,现有方法依赖于对模型权重或数值词元概率的访问,因此不适用于仅返回采样文本的接口。从采样输出中重建概率提供了一种可能的替代方案,但有限采样会产生稀疏且有噪声的估计,而在每个生成步骤重复此过程会带来大量的查询成本。我们的实证观察表明,沿着成功越狱轨迹的大规模分布变化集中在少数位置,这促使我们采用选择性控制。我们引入了\method{},一个通过仅文本续写接口进行越狱的框架,该接口允许重复采样和助手前缀续写。基于样本的分布重建将采样输出与对未观察动作的先验相结合,以获得可用的控制信号。风险门控残差控制利用不断演变的响应前缀来决定何时重建和修改分布,将采样成本集中在选定位置。推测性多词元执行通过验证和接受无需干预的草稿前缀,进一步摊销目标调用。在四个目标端点和三个基准测试中,\method{}在大多数比较中相对于基线取得了最高的平均得分。

英文摘要

Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.

发表机构

  • University of Southern California(南加州大学)
  • Adobe(奥多比)
  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

↑