AI 中文总结
研究基于LLM的内容幽默化潜在安全风险,提出HumorSafe框架评估风险传播,发现LLMs幽默化时会引入刻板印象和毒性,还提出HumorPIA攻击,揭示现有LLM安全评估在幽默化设置下的差距。
AI 中文摘要
大语言模型(LLMs)的安全防御已被广泛研究,现有方法聚焦于攻击检测和拒绝机制,但固定形式的直接拒绝策略可能引入前缀注入攻击风险。近期工作探索利用幽默作为间接拒绝机制。然而,幽默化本身是否带来安全风险尚不明。我们进行探索性研究,涉及超30000条真实智能体交互记录和45位单口喜剧演员,揭示基于LLM的内容幽默化中的实际安全问题。在此基础上,我们提出HumorSafe框架评估幽默化期间潜在安全风险传播,还提出HumorPIA提示注入攻击。实验表明HumorPIA能在保持表面安全率的同时增加毒性。我们的发现凸显了幽默化设置下现有LLM安全评估的差距。
英文摘要
Safety defenses for large language models (LLMs) have been extensively studied, with existing approaches focusing on attack detection and refusal mechanisms. Such fixed-form direct refusal strategies may introduce the risk of prefix injection attacks. Recent work has explored a new direction that leverages humor as an indirect refusal mechanism to mitigate over-refusal in jailbreak scenarios and reduce prefix injection risks. However, this approach implicitly assumes that humorous responses are safe. Whether humorization itself introduces safety risks remains unexplored. To address this issue, we conduct an exploratory study involving over 30,000 real-world agent interaction records and 45 stand-up comedians, revealing practical safety concerns in LLM-based content humorization. Motivated by these findings, we propose \textsc{HumorSafe}, a novel framework for evaluating latent safety risk propagation during humorization. \textsc{HumorSafe} enables LLMs to learn harmful humorization patterns and use them to transform benign content into humorous content with safety risks. Across five frontier LLMs, we find that LLMs can introduce stereotypes and toxicity during humorization. We further propose \textsc{HumorPIA}, a prompt injection attack that exploits latent risks in humor-based defenses. \textsc{HumorPIA} preserves the appearance of safe humorous refusal while covertly injecting harmful content, allowing latent risks to evade existing detection mechanisms. Experiments show that it increases toxicity by 3.14$\times$ while maintaining an apparent safety rate of 97.8\% even under defense settings. Our findings highlight a gap in existing LLM safety evaluations under humorized settings.