arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从预期危害性到似然性:对LLM智能体越狱的概率论重构

From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents

Juanyang Xu, Zheng Wang, Xingyu Zhao, Siddartha Khastgir, Andi Zhang

arXiv 2610.09973首次发表:更新:

发表机构

University of Macau; WMG, University of Warwick; Wuhan University(澳门大学; 华威大学WMG; 武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过概率论重构统一了最大化预期危害性与目标似然优化两种越狱方法,提出OPUR采样分布以生成高危害性目标输出并指导输入优化,实验验证其有效越狱LLM智能体。

AI 中文摘要

当LLM智能体输出的危害性可以被量化时,一个自然的越狱目标是在允许的输入修改范围内最大化预期危害性。另一种方法构造或选择有害的目标输出,并修改输入以增加其似然性。我们通过概率论重构建立了这两种方法之间的精确联系。具体来说,我们证明,关于输入的预期危害性对数的梯度,等于在危害性重加权输出分布下模型对数似然的期望输入梯度。这一恒等式为预期危害性和目标似然优化提供了统一的解释。基于这一联系,我们提出了OPUR,一种旨在生成高危害性目标输出的采样分布,并利用所得样本指导基于似然的输入优化。实验证明了该方法在越狱LLM智能体方面的有效性。

英文摘要

When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications. An alternative approach constructs or selects harmful target outputs and modifies the input to increase their likelihood. We establish a precise connection between these two approaches through a probabilistic reformulation. Specifically, we show that the gradient of the logarithm of expected harmfulness with respect to the input equals the expected input gradient of the model's log-likelihood under a harmfulness reweighted output distribution. This identity provides a unified interpretation of expected harmfulness and target likelihood optimization. Building on this connection, we propose OPUR, a sampling distribution designed to generate highly harmful target outputs and use the resulting samples to guide likelihood-based input optimization. Experiments demonstrate the effectiveness of the resulting method in jailbreaking LLM agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑