自适应对手:用于大语言模型智能体安全的多轮、多大语言模型基准测试
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
- Lambda
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出针对无记忆大语言模型防御者的自适应多轮攻击的21场景基准测试,通过观察防御者响应灵活调整攻击。介绍了不同攻击轮次成功率,汇集多个前沿攻击者大语言模型的情况,还指出不同防御者弱点差异及场景排名不一致性,并发布了相关基准测试及数据集。
AI中文摘要:
基于大语言模型的智能体在处理外部内容时,容易受到提示注入和多轮操纵的影响。大多数安全基准测试是针对评估前收集的固定攻击池来评估防御者,包括单轮或多轮。我们提出了一个针对无记忆大语言模型防御者的自适应多轮攻击的21场景基准测试:一个自主的大语言模型攻击者观察防御者的先前响应并在各轮中灵活调整,同时将每个防御者的响应视为新的交互。在保持21个场景、攻击者、防御者和结构化输出评分固定的情况下,将评分限制在攻击者的第一轮,攻击成功率为0%-1%;允许15轮自适应攻击,成功率为5.4%-14.0%。汇集三个前沿攻击者大语言模型,发现独特成功攻击数量是最佳单个攻击者的1.4-2.2倍,且生成的攻击与现有基准测试中的攻击余弦相似度低(0.02-0.14)。Claude Opus 4.6和GPT-5.4的总成功率相同(均为5.4%;95%置信区间重叠),但它们的弱点差异很大。21个场景中有13个能区分至少一对防御者,但不同场景下的排名不一致(肯德尔W=0.19)。我们发布了该基准测试,包括21个评估场景、10个公共开发场景、编排器、基线工具和多攻击者命令行界面,以及来自3×3前沿矩阵的945个记录、一个攻击重放数据集和开放竞赛决赛评分轮次的18422次gpt-oss-20b对抗。
英文摘要:
LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. We present a 21-scenario benchmark for adaptive cross-session attacks against fresh-session LLM defenders: an autonomous LLM attacker observes prior defender responses and pivots across rounds, while each defender response is evaluated as a fresh interaction. A controlled 3 x 3 attacker-defender matrix contains 945 battles. Restricting scoring to the first round yields 0-1% attack success rate (ASR); allowing 15 rounds yields 7.9-16.8%. Pooling three attacker LLMs uncovers 1.7-2.2 times as many unique successful inputs as the best single attacker, at three times the battle budget. Aggregate rates conceal opposing scenario-specific weaknesses in session-secret protection and authority handling, preserved in two higher-sample evaluations. On six scenarios, adding one provenance paragraph reduces ASR from 110/270 to 70/270, with selective effects across tasks. History and defender-state controls, together with frozen replay, characterize how the interaction protocol changes the result. A competition adds 18,422 held-out battles on a fixed gpt-oss-20b backbone and complementary benign-task evaluations. The benchmark exposes attacker and defender models, harnesses, scenarios, session state, and interaction budgets as configurable choices for systematic security evaluation.