PlanFlip:通过规划阶段提示注入攻击多智能体大语言模型系统
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
浏览论文内容
中文总结 AI 辅助
研究针对多智能体大语言模型系统规划阶段的攻击,提出PlanFlip框架包含四种攻击,通过对九个模型的3479次测试发现能力放大漏洞、同类管道盲点及推理增强模型的抗性等结果,并提出检测方法,强调异构模型多样性对系统安全的重要性。
中文摘要 AI 辅助
多智能体大语言模型系统越来越依赖规划器将目标分解为子任务序列,由下游的执行器和评论家智能体执行和审核。我们发现规划阶段是一个关键的攻击面:向规划器的上下文注入一次就能实现级联放大,同时破坏所有下游子任务。我们引入了PlanFlip框架,它包含四种规划阶段提示注入攻击——目标替换(PF-1)、优先级反转(PF-2)、上下文污染(PF-3)和角色混淆(PF-4),每种攻击都伪装成合理的工具输出以逃避关键词过滤。在对九个前沿大语言模型进行的3479次测试中,我们发现了三个结果:一是能力会放大漏洞,GPT-5的攻击成功率最高(ASR = 0.68);二是同类管道存在相关智能体盲点,如GPT-4o和Llama-3.3-70B的ASR接近0但Stealth = 1.00且StepShift > 0;三是推理增强模型能抵抗注入,如DeepSeek-R1在所有攻击中的StepShift = 0.00。我们还提出了GoalAnchorCheck(D1)和CrossAgentConsensus(D2),检测率高达1.00,在16个单元格中的15个中优于同类骨干基线。我们的关键见解是:异构模型多样性是多智能体系统的安全前提;同类骨干内的冗余无法抵御规划阶段的攻击。
英文摘要
Multi-agent LLM systems increasingly rely on a Planner to decompose goals into sub-task sequences that downstream Executor and Critic agents execute and audit. We identify the planning phase as a critical attack surface: a single injection into the Planner's context achieves cascade amplification, corrupting all downstream sub-tasks simultaneously. We introduce PlanFlip, a framework comprising four planning-phase prompt injection attacks -- GoalSubstitution (PF-1), PriorityInversion (PF-2), ContextPollution (PF-3), and RoleConfusion (PF-4) -- each disguised as plausible tool outputs to evade keyword filters. Evaluating nine frontier LLMs across 3,479 episodes, we uncover three findings: (1) capability amplifies vulnerability -- GPT-5 achieves the highest attack success rate (ASR = 0.68), contradicting the assumption that stronger models are inherently more secure; (2) homogeneous pipelines exhibit a correlated-agent blind spot -- GPT-4o and Llama-3.3-70B show ASR near 0 yet Stealth = 1.00 and StepShift > 0, with attacks restructuring plans while the same-backbone Critic reports alignment (two independent judges confirm -0.20 to -0.32 semantic deviation, r = 0.943); (3) reasoning-augmented models resist injections -- DeepSeek-R1 achieves StepShift = 0.00 across all attacks. We propose GoalAnchorCheck (D1) and CrossAgentConsensus (D2), achieving detection rates up to 1.00 and outperforming same-backbone baselines in 15 of 16 cells. Our key insight: heterogeneous model diversity is a security prerequisite for multi-agent systems; redundancy within a homogeneous backbone provides no protection against planning-phase attacks.
发表机构
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。