arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

反思与碎片:保护大语言模型免受序列马赛克攻击

Reflections and Fragments: Securing LLMs Against Sequential Mosaic Attacks

Emanuele La Malfa, Saar Cohen, Gabriele La Malfa, Mickel Liu, Christian Schroeder de Witt, Natasha Jaques, Michael J. Wooldridge

arXiv 2610.05346首次发表:更新:

发表机构

University of Oxford; Institute for Decentralized AI (IDAI); University of Washington; University College London(牛津大学; 去中心化人工智能研究所; 华盛顿大学; 伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对多轮对话中的序列马赛克攻击,提出马赛克防御理论,证明固定窗口不足,设计守望者机制实现零失败防御并保持良性有用性,实验验证了角色特定LoRA训练的有效性。

AI 中文摘要

自我对弈红队通过将攻击者和防御者角色置于零和博弈中相互对抗,提高了语言模型的安全性。然而,现实中的攻击者越来越多地使用马赛克攻击:多轮序列中的单个片段在孤立时无害,但组合起来却构成有害载荷。我们发展了一种马赛克防御理论,描述了在不牺牲有用性的前提下防止此类攻击所需的条件。我们首先证明,固定的近期提示窗口通常是不够的:安全相关信息可能出现在交互中任意远的位置。我们形式化了一个“守望者”,一种在线状态机制,将这一信息向前传递,并证明在明确假设下,它能够实现零失败防御且保持正面的良性有用性。在更强条件下,它在零失败防御者中也是最优的。然而,精确的守望者可能需要指数级多的状态,而在非结构化黑盒模型中,精确的恶意检测可能需要指数级多的查询。这些状态和查询下界本身并不意味着学习困难:状态下界背后的构造可以从标记示例中高效学习,而在受限访问下,认证最坏情况安全性可能需要显著更多的信息。我们还表明,仅靠自我对弈均衡并不能保证有用性,这促使了一种约束公式化,即在零失败防御者中最大化最坏情况下的良性有用性。实验上,通过多轮自我对弈在冻结的LLM上训练角色特定的攻击者和防御者LoRA适配器,增强了两个角色:攻击者在引发有害响应方面变得更有效,而防御者对攻击变得更鲁棒,且在未见过的攻击目标上也观察到了改进。

英文摘要

Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isolation yet assemble into a harmful payload. We develop a theory of mosaic defense that characterizes what is required to prevent such attacks without sacrificing helpfulness. We first show that no fixed bounded window of recent prompts is sufficient in general: safety-relevant information may occur arbitrarily far back in the interaction. We formalize a watchman, an online state mechanism that carries this information forward, and show that under explicit assumptions it enables zero-failure defense with positive benign helpfulness. Under stronger conditions, it is also optimal among zero-failure defenders. An exact watchman may nevertheless require exponentially many states, while exact maliciousness detection can require exponentially many queries in an unstructured black-box model. These state and query lower bounds do not by themselves imply hard learning: the construction underlying the state lower bound is efficiently learnable from labeled examples, whereas certifying worst-case safety can require substantially more information under restricted access. We also show that self-play equilibrium alone does not certify usefulness, motivating a constrained formulation that maximizes worst-case benign helpfulness among zero-failure defenders. Empirically, training role-specific attacker and defender LoRA adapters over frozen LLMs via multi-turn self-play strengthens both roles: attackers become more effective at eliciting harmful responses, while defenders become more robust to attack, with improvements also observed on unseen attack objectives.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑