发表机构
Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ActGov提出运行时策略约束验证框架,逐行动校验LLM工具调用,通过SMT反例检查构建策略,在AgentDojo和AgentDyn上显著降低间接提示注入成功率,同时保持任务效用。
AI 中文摘要
大型语言模型(LLM)智能体日益通过外部工具执行长时程工作流,这使得不受信任的输出能够影响后续行动并超越用户授权。现有防御措施隔离注入内容或通过预定义计划和静态策略约束执行,但这些方法在动态工作流下显得脆弱,并且在可扩展的工具生态系统中扩展性不佳。在本工作中,我们提出了ActGov,一个运行时强制执行框架,它在每个LLM提议的工具行动产生外部影响之前对其进行验证。基于统一的授权、行动、运行时上下文和安全约束语义模型,ActGov-Policy组件从工具规范、良性任务和观察到的失败轨迹中迭代构建策略集,每次更新都通过基于SMT的反例检查进行验证。在运行时,ActGov-Runtime将每个工具调用抽象为有限的策略记录,并且仅当该调用保持在任务范围内的授权边界内并满足所有适用策略时才允许其执行。这种逐行动强制执行在长时程、动态分支工作流中保持了授权。我们在AgentDojo和AgentDyn基准上,跨多个模型和攻击配置对ActGov进行了评估。结果表明,ActGov持续降低了间接提示注入攻击的成功率,同时保持了任务效用,显著优于现有防御措施。这些结果证明,ActGov能够在动态智能体执行中强制执行细粒度授权,而无需依赖底层LLM来正确识别恶意指令。
英文摘要
Large language model (LLM) agents increasingly execute long-horizon workflows through external tools, allowing untrusted outputs to influence subsequent actions and exceed user authorization. Existing defenses isolate injected content or constrain execution with predefined plans and static policies, but these approaches are brittle under dynamic workflows and scale poorly across extensible tool ecosystems. In this work, we present ActGov, a runtime enforcement framework that validates each LLM-proposed tool action before it causes external effects. Built on a unified semantic model of authorization, actions, runtime context, and security constraints, the ActGov-Policy component iteratively constructs a policy set from tool specifications, benign tasks, and observed failure traces, with each update verified through SMT-based counterexample checking. At runtime, ActGov-Runtime abstracts each tool call into finite policy records and permits it only if it remains within the task-scoped authorization boundary and satisfies all applicable policies. This per-action enforcement preserves authorization throughout long-horizon, dynamically branching workflows. We evaluate ActGov on the AgentDojo and AgentDyn benchmarks across multiple models and attack configurations. It shows that ActGov consistently reduces the success rate of indirect prompt-injection attacks while preserving task utility, significantly outperforming existing defenses. These results demonstrate that ActGov can enforce fine-grained authorization over dynamic agent executions without relying on the underlying LLM to correctly identify malicious instructions.