发表机构
Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究发现,针对跨不可链接身份的分解攻击,现有有状态防御策略在攻击者可重试学习时无法有效阻止攻击,需引入身份链接等额外分组相关机制。
AI 中文摘要
大多数大语言模型(LLM)服务使用仅判断当前请求的无状态防御来拒绝有害任务。分解攻击利用这一局限性,将有害任务拆分为多个单独允许的请求并组合其答案。因此,防御此类攻击需要考虑请求关联的有状态监控器,若该监控器能对同一攻击者任务的所有请求进行分组,即可阻止攻击。然而,攻击者可使用不可链接的身份并在其他地方组合答案,从而留下可靠的分组信号。本文研究在该设置下是否仍能阻止分解攻击。对于固定且无重试的攻击策略,本文证明可实现的安全性与效用的权衡完全取决于如何对具有相同能力的良性请求进行分组:持久且可识别的分组允许实用的防御;而全新且不可区分的分组则不行。当攻击者可重试并从允许/阻止决策中学习时,该实用操作点消失:反馈揭示了哪些请求通过,但未揭示阻止是否正确。在91项可执行任务和11393个能力匹配的良性请求上进行的实验支持了这些结果:在对这些请求设置1%的拒绝上限、对无关背景流量设置0.5%上限的条件下,包括拥有精确请求-操作映射的特权策略在内的所有10种测试策略,要么无法阻止攻击,要么超出预算;在防御未见过的任务族上,攻击成功率在一次尝试后至少为99%,两次尝试后达100%。因此,有效的防御需要与分组相关的额外证据或机制,例如可靠的身份链接、新身份的成本或对答案使用的控制。
英文摘要
Most large language model services use stateless defenses, which judge only the current request, to refuse harmful tasks. Decomposition attacks exploit this limitation by splitting a harmful task into individually permissible requests and combining their answers. Defending against them therefore requires a stateful monitor that considers requests together. If it can group all requests for one attacker task, it can stop the attack. However, attackers can use unlinkable identities and combine answers elsewhere, leaving no reliable grouping signal. We ask whether decomposition attacks can still be stopped under this setting. For a fixed attack strategy without retries, we prove that the achievable security and utility tradeoff depends entirely on how benign requests for the same capabilities are grouped. Persistent, recognizable groups permit a useful defense; fresh, indistinguishable groups do not. When attackers can retry and learn from Allow/Block decisions, this useful operating point disappears: the feedback reveals what passes but not whether a block was correct. Experiments on 91 executable tasks and 11,393 capability-matched benign requests support these results. Under a 1% denial cap for these requests and a 0.5% cap for unrelated background traffic, all ten tested policies, including one privileged policy with an exact request-to-operation map, either fail to stop attacks or exceed the budget. On defense-unseen task families, attack success is at least 99% after one attempt and 100% after two. Effective defenses therefore require additional evidence or mechanisms tied to grouping, such as reliable identity linkage, costs for fresh identities, or control over answer use.