发表机构
University of Arizona(亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出信息论框架分析组合式意图隐藏越狱,通过先验-后验匹配选择辅助任务,证明精确匹配计算困难并给出最优解法,实验显示组合查询可提升越狱效果但增大捆绑包会降低目标保留。
AI 中文摘要
近期研究表明,大型语言模型(LLMs)可能易受越狱攻击的影响,此类攻击通过将有害意图与良性任务组合来掩盖其真实目的。一个单独提出会被拒绝的有害请求,当嵌入到更大、看似良性的查询中时,可能会引发不同的响应。我们从信息论视角研究这些组合式意图隐藏越狱。我们的公式将每个任务与一个被判定为有害的估计概率相关联:整个任务集合上的平均值定义了有害意图的先验概率,而包含目标的选定捆绑包上的平均值则定义了后验概率。选择辅助任务使这些平均值一致(我们称之为先验-后验匹配),即使有害目标仍保留在捆绑包中,估计的意图也保持不变。我们研究了两种设置,区别在于查询构建是否属于优化的一部分。在查询无关设置中,任务的选择不考虑它们在最终查询中的表达方式。我们证明了在捆绑包大小约束下精确的先验-后验匹配在计算上是困难的,推导了分数权重的最优注水解法,并刻画了满足规定安全阈值的最小捆绑包。在查询相关设置中,任务选择与查询构建联合考虑,意图隐藏和目标保留在生成的查询上进行评估。我们评估了跨捆绑包大小、查询生成器和多个开源模型的越狱有效性和目标行为保留情况。结果表明,在所评估的搜索预算下,组合查询可以引发超出直接请求基线之外的目标行为,同时揭示了一个权衡:随着捆绑包大小的增加,几个模型的响应级目标保留往往下降。
英文摘要
Recent work has shown that large language models (LLMs) can be vulnerable to jailbreak attacks in which harmful intent is obscured through composition with benign tasks. A harmful request refused in isolation may elicit a different response when embedded within a larger, seemingly benign query. We study these compositional intent-hiding jailbreaks from an information-theoretic perspective. Our formulation associates each task with an estimated probability of being judged harmful: the average over the full task collection defines the prior probability of harmful intent, while the average over a selected bundle containing the target defines the posterior. Selecting auxiliary tasks so that these averages agree, which we call prior-posterior matching, leaves the estimated intent unchanged even though the harmful target remains in the bundle. We study two settings that differ in whether query construction is part of the optimization. In the query-independent setting, tasks are selected without regard to how they will be expressed in the final query. We show that exact prior-posterior matching under a bundle-size constraint is computationally hard, derive an optimal water-filling solution for fractional weights, and characterize the smallest bundle satisfying a prescribed safety threshold. In the query-dependent setting, task selection and query construction are considered jointly, and intent concealment and target preservation are evaluated on the resulting query. We evaluate jailbreak effectiveness and preservation of the target behavior across bundle sizes, query generators, and several open-source models. These results show that compositional queries can elicit target behaviors beyond the direct-request baseline under the evaluated search budgets, while revealing a trade-off: as bundle size increases, response-level target preservation tends to decrease for several models.