arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ASCII攻击:将有害请求重新语境化为大型语言模型中的艺术评论

ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models

Da Cheng Gu, Yifei Dong, Xinghao Yang, Yongshun Gong, Wei Liu

arXiv 2609.02215首次发表:更新:

发表机构

Faculty of Engineering and Information Technology, University of Technology Sydney; China University of Petroleum; Shandong University(悉尼科技大学工程与信息技术学院; 中国石油大学; 山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出ASCII攻击,通过将有害请求嵌入ASCII艺术并以艺术评论形式呈现,可绕过大型语言模型的安全对齐,在11个模型、8个主题上使62%的框架化提示被判定有害,最易受攻击模型成功率达93%。

AI 中文摘要

安全对齐训练大型语言模型以拒绝明确表述的有害请求,但该训练主要针对表层形式。仅对相同操作内容进行重新语境化、改变模型解读方式的请求,因此仅能被弱覆盖。ASCII攻击便是此类重新语境化手段之一,它是单轮且黑盒的:仅一条消息,无法访问模型内部结构。该攻击将完全清晰的有害请求嵌入ASCII艺术字符中,呈现为艺术品,并请求反馈。与ArtPrompt不同,它不隐藏任何内容:请求保持可读。其回复以艺术评论形式撰写,可包含明确请求会被拒绝的操作细节。每个被框架化的提示都配有直接问题对照,因此对比结果与主题、模型和解码变化无关。该对比识别出的是一个捆绑的表层,而非孤立通道。在11个模型和8个有害主题上,有害感知分类器判定62%的框架化提示有害,而对照组的这一比例为42%。在最易受攻击的模型上,框架化提示的成功率达93%。单次查询在5个有害判定器中有4个下,与已发表的单次查询攻击相当或更优。该效果更多取决于模型而非主题,且不会随模型规模增大而减弱。近三分之二的框架化行中,至少有一个判定器与多数意见不同,这本身是一项测量有效性发现。该模式与不匹配的泛化一致。

英文摘要

Safety alignment trains large language models to refuse harmful requests stated plainly, but that training is applied mostly to surface form. Requests that only recontextualise the same operational content, changing how the model reads it, are therefore only weakly covered. The ASCII Attack is one such recontextualisation. It is single-turn and black-box: one message, with no access to model internals. It embeds a fully legible harmful request in ASCIl-art characters, presents it as artwork, and asks for feedback. Unlike ArtPrompt, it hides nothing: the request stays readable. The reply is written as artistic critique and can contain operational detail that a plain request would have been refused for. Every framed prompt is paired with a direct-question control, so the contrast is isolated from topic, model and decoding variation. The contrast identifies a bundled surface, not one isolated channel. Across eleven models and eight harm topics, a harm-aware classifier judges 62% of framed prompts harmful against 42% of controls. On the most susceptible model the framed prompt succeeds 93% of the time. A single query matches or exceeds published single-query attacks under four of five harm judges. The effect tracks the model more than the topic and does not diminish with scale. At least one judge dissents from the panel majority on nearly two-thirds of framed rows, which is itself a measurement-validity finding. That pattern is consistent with mismatched generalisation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑