arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27373cs.CRcs.AI

RoguePrompt:用于自重构以规避大型语言模型(LLM)审核的双层编码

RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation

Benyamin Tafreshian, Prathamesh Dhake

首次发表
浏览论文内容

中文总结 AI 辅助

RoguePrompt是一种采用双层编码(维吉尼亚密码+ROT13)的LLM越狱流程,在黑盒威胁模型下对313个被拒提示词测试,实现93.93%的过滤绕过率,提供多阶段越狱失效的阶段级证据。

中文摘要 AI 辅助

大型语言模型(LLM)正日益融入主流开发平台和日常技术工作流,通常部署在审核与安全控制之后。尽管存在这些控制,防止基于提示词的政策规避仍具挑战性,攻击者持续通过精心制作提示词来规避已实施的安全机制,从而对LLM进行“越狱”。现有研究已证实,密码介导交互、嵌入代码的解密、提示词分解与重构、分层自定义加密是可行的攻击原语。然而,已报道的评估通常将可见接受度、隐藏请求的成功恢复及后续执行合并为单一的攻击成功结果,这使得关于多阶段提示词转换攻击在可观测的黑盒交互中何处失效的证据有限。本文提出RoguePrompt,这是一种越狱流程,它将禁止的提示词进行划分,并应用维吉尼亚密码(Vigenere)和ROT13两种嵌套编码,同时结合自然语言重构指令。RoguePrompt在黑盒威胁模型下开发与评估,仅通过API或用户界面访问托管模型,测试对象为313个真实世界中被强硬拒绝的提示词。成功的衡量标准包括审核绕过、指令重构及执行,当相关阶段超出其自动标准时即判定成功。RoguePrompt在过滤器绕过上达到93.93%的平均成功率,在重构上为79.02%,在执行上为70.18%。这些结果证明了分层提示词编码的有效性,同时提供了多阶段越狱在审核绕过、指令重构及执行过程中何处失效的阶段级证据。

英文摘要

Large language models (LLMs) are becoming increasingly integrated into mainstream development platforms and daily technological workflows, typically behind moderation and safety controls. Despite these controls, preventing prompt-based policy evasion remains challenging, and adversaries continue to "jailbreak" LLMs by crafting prompts that circumvent implemented safety mechanisms. Prior work has established cipher-mediated interaction, code-embedded decryption, prompt decomposition and reconstruction, and layered custom encryption as viable attack primitives. However, reported evaluations generally collapse visible acceptance, successful recovery of the concealed request, and subsequent execution into an aggregate attack-success outcome. This leaves limited evidence about where multistage prompt-transformation attacks fail within an observable black-box interaction. This paper introduces RoguePrompt, a jailbreak pipeline that partitions a forbidden prompt and applies two nested encodings, Vigenere followed by ROT13, along with natural-language reconstruction instructions. RoguePrompt was developed and evaluated under a black-box threat model, with only API or user-interface access to the hosted models, and was tested on 313 real-world, hard-rejected prompts. Success was measured in terms of moderation bypass, instruction reconstruction, and execution when the relevant stage exceeded its automated criterion. RoguePrompt achieved average rates of 93.93% for filter bypass, 79.02% for reconstruction, and 70.18% for execution. These results demonstrate the effectiveness of layered prompt encoding while providing stage-level evidence of where multistage jailbreaks fail during moderation bypass, instruction reconstruction, and execution.

发表机构

  • Boston University(波士顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑