用于大语言模型代理中污点限制的代理权限策略代数
APPA: Recoverable Information-Flow Control for Real-World LLM Agents
- Archestra AI
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究大语言模型代理处理混合机密数据时的安全风险,提出APPA框架,通过引擎管理上下文分支等解决可用性瓶颈,经双幺半群模型证明相关特性,实验表明其能抑制数据泄露并恢复部分被污点跟踪损失的效用。
AI中文摘要:
处理混合机密性数据的自主大语言模型代理面临来自提示注入攻击和推理错误的严重安全风险。虽然动态信息流控制提供了结构安全保证,但传统的污点跟踪在读取未经审查的数据时会永久污染代理的上下文环境,严重限制了下游效用。我们提出了APPA(代理权限策略代数),这是一个通过引擎管理的上下文分支和预期获取执行来解决此可用性瓶颈的信息流控制框架。在数据获取之前,APPA会预期评估标签下降和缺失的前提条件,生成可操作的补救计划(授权、接受)。为了在不污染主上下文的情况下检查未经审查的数据,会生成一个由标签种子化的子轨迹,在本地吸收标签下降,并允许可信的净化器向未改变的父轨迹返回有界导数。由安全标签和共享事件日志上的双幺半群模型管理,我们正式证明了父标签保留和合并限制。最后,我们在四个模型的多轮工具链基准测试中评估了APPA:它抑制了数据泄露(攻击成功率从31%-50%降至0%-7%),并且在四个模型中的三个模型上,分支恢复了仅污点跟踪所损失的大量效用。
英文摘要:
LLM agents deployed in practical workflows routinely mix private context, untrusted tool and web outputs, and external side effects. While information-flow control (IFC) provides structural defenses against prompt injection, data exfiltration, and confused-deputy attacks, conventional IFC relies on monotone taint tracking that either over-blocks benign operations or permanently strands downstream execution once an agent ingests unvetted data. We present APPA (Agentic Permissions Policy Algebra), which turns agent IFC from an abort-only barrier into a policy-governed recovery system. APPA enforces a dual-phase reference monitor at tool dispatch and protocol gateways (e.g., Model Context Protocol): before tool execution, it prospectively evaluates composite label restrictions and workflow history; upon completion, it validates realized outputs before context admission. For incremental rollout across unannotated tools, APPA incorporates gradual security typing with bounded cast resolution. To inspect untrusted data without poisoning primary agent context, APPA introduces on-demand trajectory confinement: disposable child branches absorb taint locally and exit through shape-bounded channels (attest-schema) with exact parent-label and transcript preservation, avoiding permanently partitioned multi-agent infrastructure. We prove core safety invariants: no-laundering gradual resolution, branch boundary isolation, and recovery containment against prompt-injected models. Across 6,600 controlled benchmark episodes spanning OWASP AgentThreatBench and enterprise workflows (Bench-Corp), APPA sustains 64.2-91% utility with zero observed attacks across 1,320 guarded episodes, establishing a practical defense for deployed tool-using agents.