ContainmentBench:基于跟踪的工具使用大型语言模型代理注入后遏制评估
ContainmentBench: Trace-Based Evaluation of Post-Exposure Containment in Tool-Using LLM Agents
浏览论文内容
中文总结 AI 辅助
研究针对工具使用的大型语言模型代理注入后遏制评估问题,引入ContainmentBench基准,分别测量端点策略合规性等。通过对Qwen2.5 - 7B - Instruct研究发现,终端策略标签不充分,评估应分别报告多方面证据,还给出不同策略的工作流程完成率等结果。
中文摘要 AI 辅助
使用工具的大型语言模型代理处理不受信任的内容,维护内存,跨代理委托,并调用有副作用的工具。现有的提示注入评估通常用终端攻击或策略结果来总结安全性,但相同的端点可能隐藏不同的暴露后跟踪和授权效用的不同损失。我们引入了ContainmentBench,这是一个基于跟踪的沙盒基准,分别测量基准定义的端点策略合规性、记录的传播、恢复检测和授权的结构化操作完成情况。在对Qwen2.5 - 7B - Instruct进行的预先指定的17640次展开研究中,比较仅污点和意图感知执行的所有600对匹配的活动 - 受污染对具有相同的零伤害结果,但73.5%在记录的轨迹或效用方面存在差异。仅污点执行仅完成0.1642的授权污点工作流程;可信账本策略将完成率提高到0.8567,而在相同的观察到的端点策略结果下,强大的工具边界基线达到0.9233。我们还发现,汇总的记录传播排名会随着证据阶段组成和分母选择而变化。这些结果表明,终端策略标签对于操作后的暴露后遏制不是一个充分的统计量;评估应分别报告端点、阶段分层轨迹和效用证据,并且仅在相应控制有效的情况下才将恢复证据用于比较性声明。全面研究是合成的且是单模型的;政策案例还假设了一个正确的结构化授权账本。
英文摘要
Tool-using large language model (LLM) agents read untrusted content, maintain memory, delegate tasks, and invoke tools with external side effects. Terminal attack-success or policy-violation rates do not show what happens between exposure and commit or whether a defense also suppresses authorized actions. We introduce ContainmentBench, a sandboxed benchmark comprising a 504-scenario specification dataset, a shared rollout-trace schema, and stage-scoped metrics for endpoint violations, logged propagation, and explicitly authorized taint-exposed proposals that commit. The main Qwen2.5-7B-Instruct study evaluates seven policy conditions and five seeds, yielding a 17,640-record trace corpus. Across 600 matched active-tainted rollout pairs, no committed policy violation was observed under either taint-only or intent-ledger enforcement. Their execution records nevertheless differed: 441 pairs (73.5%) had different values in a shared 12-field trace summary that includes commit-related diagnostics, and the mean authorized proposal-commit score was 0.164 under taint-only enforcement and 0.857 under intent-ledger enforcement, compared with 0.923 under tool-boundary enforcement. Logged-propagation rankings changed with stage selection and normalization. In a limited set of custom AgentDojo-native workflows, committed violations were observed without defense and were not observed under either evaluated defense. A separate 6,048-rollout Mistral/common-JSON model-interface configuration retained the v1-to-v2 proposal-commit improvement, but committed violations were observed under intent-ledger v2. Equal terminal outcomes do not imply equal containment. The evaluation uses synthetic workflows. The intended intent-ledger mechanism assumes schema-aligned authorization metadata; one public-status task family violates this assumption and is analyzed separately.