arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相同结果,不同证据:大语言模型安全评估中的意图恢复

Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation

Haitong Jiang, Chunlin Liu, Sihan Tang, Chan Wu, Xiaoqing Su, Yuhong Feng

arXiv 2610.11766首次发表:更新:

发表机构

College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机与软件学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型安全评估仅用攻击成功率(ASR)的不足,提出用操作理解率(UR)配对评估,发现非有害结果信息价值不均,建议联合报告意图恢复与ASR。

AI 中文摘要

大语言模型的安全评估通常用攻击成功率(ASR)来总结有害输出行为。然而,相同的非有害结果可能源于截然不同的原因:模型可能恢复有害任务并拒绝执行,可能无法恢复该任务,也可能完全响应其他内容。在意图模糊的提示下,区分这些情况尤为重要,此时低ASR无法表明被评估任务是否实际被调用。为明确区分,我们将ASR与操作理解率(UR)配对,UR用于衡量响应是否既识别了被评估任务,又将其视为待回答的任务。在不同交互界面中,这种配对视角揭示了ASR所隐藏的显著差异:相似的ASR值可能对应截然不同的恢复率。受控英文重构实验显示,随着压缩提示变得更明确,恢复率持续提升,而ASR并未遵循相同模式。互补对比来自FormalLogic数据集,高恢复率仍可能伴随频繁的有害协助。这些结果共同表明,非有害结果对模型安全性的信息价值并不均等,这促使在大语言模型安全评估中联合报告意图恢复和ASR。代码与实验输入可在该https URL获取。

英文摘要

Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing these cases becomes especially important under intent-obscuring prompts, where a low ASR does not reveal whether the evaluated task was actually engaged. To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered. Across interfaces, this paired view reveals substantial variation hidden by ASR: similar ASR values can correspond to sharply different recovery rates. Controlled English reconstructions show that recovery consistently improves as compressed prompts become more explicit, whereas ASR does not follow the same pattern. A complementary contrast comes from FormalLogic, where high recovery can still coincide with frequent harmful assistance. Together, these results show that non-harmful outcomes are not equally informative about model safety, motivating the joint reporting of intent recovery and ASR in LLM safety evaluation. Code and experiment inputs are available at https://github.com/kevinjiang0121-cyber/IRIS.

Comments12 pages, 2 figures. Code and experiment inputs: https://github.com/kevinjiang0121-cyber/IRIS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑