妥协并非后果:使用配对重放评估LLM智能体中的任务范围授权
Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replay
浏览论文内容
中文总结 AI 辅助
本研究通过配对重放测试平台评估任务范围授权对LLM智能体工具执行的控制效果,发现范围条件可完全遏制有害执行,支持遏制而非预防提示注入。
中文摘要 AI 辅助
使用工具的模型即使在凭据有效时也可能遵循恶意指令。我们研究任务范围授权是否能控制由此产生的工具执行。我们的配对重放测试平台对模型请求进行一次采样,并将相同的操作、资源和参数提交给广泛持有者、范围JWT、发送者约束和开放策略代理条件。冻结的主实验使用四个工具域和五个本地模型配置中的128个场景。在有效的受攻击后暴露决策中,广泛持有者有害执行率在模型间为8.9%至37.8%;所有三个范围条件均为零。策略改变执行,而非冻结的模型决策。这些结果支持在研究者提供的任务权限下,对所测试的跨操作和跨资源后果进行遏制,而非防止提示注入。一个有界的六任务AgentDojo扩展保留了原生评分,并记录了11/24的注入广泛策略攻击标志,而范围标志为0/24,效用为5/24和6/24。低暴露、无效决策和有目的的任务选择限制了该比较。主要实验衡量安全继续而非最终答案的正确性。一项经验库新鲜采样比较发现,当范围结果为恒定零时,配对没有估计优势。贡献是对妥协、执行和继续的受控测量,并明确说明每个观察所确立的界限。
英文摘要
A tool-using model can follow a malicious instruction even when its credentials are valid. We study whether task-scoped authorization contains the resulting tool execution. Our paired-replay testbed samples a model request once and submits the same action, resource, and arguments to broad bearer, scoped JWT, sender-constrained, and Open Policy Agent conditions. The frozen main experiment uses 128 scenarios across four tool domains and five local model configurations. Among valid attacked post-exposure decisions, broad-bearer harmful execution ranges from 8.9% to 37.8% across models; all three scoped conditions record zero. The policy changes execution, not the frozen model decision. These results support containment of the tested cross-action and cross-resource consequences under researcher-supplied task authority, not prevention of prompt injection. A bounded six-task AgentDojo extension preserves native scoring and records 11/24 injected broad-policy attack flags versus 0/24 scoped flags, with utility of 5/24 and 6/24. Low exposure, invalid decisions, and purposive task selection limit that comparison. The primary experiment measures safe continuation rather than final-answer correctness. An empirical-bank fresh-sampling comparison finds no estimation advantage from pairing when scoped outcomes are constant zero. The contribution is a controlled measurement of compromise, enforcement, and continuation, with explicit limits on what each observation establishes.