arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02681cs.CR

批量提示中的安全性?理解并缓解批量提示中的安全故障

Safety in Batches? Understanding and Mitigating Safety Failures in Batch Prompting

  • Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Kihyun Kim, Hee-Seon Kim, Wonjun Lee, Changick Kim

AI总结:

该研究发现批量提示存在独特安全故障模式,可作为黑盒攻击实现高成功率,提出感知批量的偏好优化可缓解该漏洞,为大语言模型安全对齐提供了新方向。

AI中文摘要:

批量提示是大语言模型的实用推理策略,但其安全影响仍未得到充分探索。我们表明,批量提示在效用方面的成功并未延伸到安全领域:一个单独时会被可靠拒绝的有害问题,嵌入在一组良性问题中时会引发有害响应。我们将此确定为一种独特的安全故障模式,无法归为上下文学习或长上下文效应等已知漏洞,并从两个互补角度分析其原因:对齐信号减弱和拒绝信号稀释。在广泛使用的开源模型和前沿商业模型中,批量提示作为一种简单的黑盒攻击始终实现较高的攻击成功率。我们进一步表明,感知批量的偏好优化可有效缓解该漏洞。这些发现凸显了当前安全对齐中的一个盲点,并指出感知批量的对齐是实现稳健部署的必要步骤。

英文摘要:

Batch prompting is a practical inference strategy for large language models, but its safety implications remain underexplored. We show that the success of batch prompting for utility does not extend to safety: a harmful question that is reliably refused in isolation can elicit a harmful response when embedded in a batch of benign questions. We identify this as a distinct safety failure mode -- not reducible to known vulnerabilities such as in-context learning or long-context effects -- and analyze its causes from two complementary perspectives: alignment signal weakening and refusal signal dilution. Across widely used open-source and frontier commercial models, batch prompting consistently achieves high attack success rates as a simple black-box attack. We further show that batch-aware preference optimization effectively mitigates the vulnerability. These findings highlight a blind spot in current safety alignment and point to batch-aware alignment as a necessary step toward robust deployment. Code is available at https://github.com/96kihyun/batch_jailbreak

补充信息

↑