叙事包装如何影响大语言模型的拒绝:跨语言基准与防御方法
How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
浏览论文内容
中文总结 AI 辅助
该研究针对大语言模型在叙事包装下的拒绝漏洞,构建跨语言基准GUISE,提出AXIS防御方法,在多款模型上实现了最优的安全与可用性平衡。
中文摘要 AI 辅助
安全对齐的语言模型通常会直接拒绝明确表述的有害请求,但会回答角色扮演或叙事包装内的相同请求。我们在不同语言和语域中测量这种漏洞:Qwen3-1.7B在英文上的攻击成功率已达89.4%,现代汉语上为93.0%,古汉语上则达到95.7%。我们构建了GUISE,这是一个用于系统研究该漏洞的基准,包含英文、现代汉语和古汉语的平行请求、匹配的有害与良性对、用于评估的预留包装类型,以及将“警告后回答”响应计为攻击成功的更严格标准。表征分析显示,语言和语域仅使有害请求的表征略微偏离模型的拒绝方向,而叙事包装则使其偏离得远得多。我们提出AXIS,它将偏好优化与旋转目标相结合,使有害请求的表征与拒绝方向对齐,同时结合承诺目标,训练模型完全拒绝而非产生“警告后回答”响应。在Qwen3-1.7B、Qwen3-4B和GLM-4-9B上,AXIS在对比方法中实现了最高的安全与可用性综合得分。
英文摘要
Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther away. We propose AXIS, which combines preference optimisation with a rotation objective that aligns harmful-request representations with the refusal direction and a commitment objective that trains the model to refuse completely rather than produce a warn-then-answer response. Across Qwen3-1.7B, Qwen3-4B and GLM-4-9B, AXIS achieves the highest combined safety and usability score among the compared methods.
发表机构
- Florida State University(佛罗里达州立大学)
- Texas Christian University(得克萨斯基督教大学)
- University of Pennsylvania(宾夕法尼亚大学)
- Texas Tech University(得克萨斯理工大学)
机构由 AI 辅助整理,请以论文原文为准。