arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测量错误的指标:内部有害性评分反排成功的越狱攻击

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

Mingyu Luo, Ming Deng, Zilang Qiu, Yiming Cheng, Ci Tao, Xue Tan, Sijin Sun, Yangfu Li, Ping Chen, Jun Dai, Xiaoyan Sun

arXiv 2608.09624首次发表:更新:

发表机构

College of Computer Science and Artificial Intelligence, Fudan University; School of Computer Engineering and Science, Shanghai University; Beijing Normal University; Institute of Advanced Intelligence and Computing, A*STAR; School of Communication and Electronic Engineering, East China Normal University; Institute of Big Data, Fudan University; Department of Computer Science, Worcester Polytechnic Institute(复旦大学计算机科学与人工智能学院; 上海大学计算机工程与科学学院; 北京师范大学; 新加坡科技研究局高级智能与计算研究所; 华东师范大学通信与电子工程学院; 复旦大学大数据研究院; 伍斯特理工学院计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文审计内部安全评分的推论逻辑,提出主动注意力探测方法,发现包装操作使攻击更危险但评分误判,反转在多模型等场景持续存在,揭示内部评分的测量偏差问题。

AI 中文摘要

内部安全评分在生成任何文本前判断提示词,其有效性通过区分有害提示词与良性提示词的能力验证,这种区分被视为该评分也能捕获成功攻击的证据。有害意图是提示词的属性,而越狱成功是特定目标模型、解码策略和评判器共同产生的结果。针对测量错误量的评分调整的过滤器,会将误判预算浪费在本会失败的攻击上。本文对该推论进行审计:基于注意力的测量通常从依赖提示词的位置读取,因此包装器会同时改变被判断的内容和信号提取位置,为此我们引入主动注意力探测(Active Attention Probing),提供固定的、与内容无关的测量坐标。我们为每个基础目标搭配普通版和包装版,从目标模型生成真实补全内容。在Llama模型上,包装操作使有害生成率从0.05升至0.27,同时有害意图的AUROC从0.936降至0.803,即攻击变得更危险,而提示词在评分看来更安全;在包装后的有害提示词中,结果AUROC为0.220,这使得成功的攻击排在失败攻击之下。稀有标记、被动及检测器衍生的通道在相同匹配设计上重现了该反转,且该反转在3种目标模型、7种攻击家族和2个独立评判器中持续存在,分布偏移会在校准和阈值转移降级前先降级排序。

英文摘要

Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt. Jailbreak success is an outcome produced later by a particular target model, decoding policy, and judge. A filter tuned on a score that measures the wrong quantity spends its false positive budget on attacks that would have failed anyway. In this paper we audit that inference. Attention based measurements are usually read from prompt dependent locations, so a wrapper changes both the content being judged and the place the signal is taken from. We therefore introduce Active Attention Probing, which supplies a fixed content independent measurement coordinate. We pair every base goal with a plain and a wrapped version and generate real completions from the target models. On Llama, wrapping raises harmful generation from 0.05 to 0.27 while harmful intent AUROC falls from 0.936 to 0.803, so the attacks grow more dangerous while the prompts look safer to the score. Among wrapped harmful prompts the outcome AUROC is 0.220, which places the attacks that succeeded below the attacks that failed. Rare token, passive, and detector derived channels reproduce the reversal on the same matched design, and the reversal itself persists across three target models, seven attack families, and two independent judges. Distribution shift then degrades calibration and threshold transfer before it degrades ranking.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑