AgentFlow:一种以流为中心的策略语言与框架,用于保障大语言模型智能体系统的安全
AgentFlow: A Flow-Centric Policy Language and Framework for Securing LLM Agent Systems
浏览论文内容
中文总结 AI 辅助
本文提出以流为中心的AgentFlow策略语言与框架,通过运行时监控与SMT验证保障LLM智能体系统安全,在多基准测试中显著降低入侵率并提升效用。
中文摘要 AI 辅助
大语言模型(LLM)智能体日益频繁地读取不可信内容、调用外部工具、访问私有数据,并将工作委派给其他智能体。危害往往并非源于单一的不安全操作,而是敏感数据在一系列看似合理的步骤间的流动所引发。本文提出AgentFlow,一种以流为中心的策略语言与运行时执行模型,用于指定智能体系统中数据的允许流向。策略基于带标签的运行时边定义,约束哪些工具可接收敏感字段、哪些接收端可接收释放的数据,以及哪些权限可跨越委派边界。该语言支持流规则与路径规则、任务范围的权限、受控释放,以及有状态的污点语义。一个运行时参考监视器调解智能体操作,一个基于有界SMT的验证器针对结构化策略片段检查安全属性。我们在多个智能体基准上评估AgentFlow:在原型系统中,7个安全属性每个的验证耗时均在0.5秒以内,且验证器捕获了研究中所有植入的不安全策略变体;在4个套件共949个AgentDojo注入案例中,AgentFlow将已确认的入侵率从33.0%降至0.0%,同时将综合效用从46.7%提升至63.3%;在200个案例的AgentDyn日常生活基准中,其将已确认的入侵率从73.5%降至0.0%,同时保留接近基线的效用(从44.5%变为43.5%);对ASB、InjecAgent、BIPIA、AgentHarm和MCPTox重放的广度检查显示,配置的策略可阻断基准指定的策略可见攻击者流,在ASB的直接提示注入工具中,攻击成功率为0/1200。这些结果为初步结果,仅适用于建模的策略可见智能体行为及所评估的基准。
英文摘要
LLM agents increasingly read untrusted content, invoke external tools, access private data, and delegate work to other agents. Harm often arises not from a single unsafe action but from the flow of sensitive data across a sequence of otherwise plausible steps. We present AgentFlow, a flow-centric policy language and runtime enforcement model for specifying where data may travel in agent systems. Policies are defined over labeled runtime edges and constrain which tools may receive sensitive fields, which sinks may receive released data, and what authority may cross delegation boundaries. The language supports flow and path rules, task-scoped capabilities, controlled release, and stateful taint semantics. A runtime reference monitor mediates agent actions, and a bounded SMT-based verifier checks safety properties for a structured policy fragment. We evaluate AgentFlow on multiple agent benchmarks. In our prototype, seven safety properties verify in under 0.5 seconds each, and the verifier catches all seeded unsafe policy variants in our study. On 949 AgentDojo injected cases across four suites, AgentFlow reduces confirmed compromise from 33.0\% to 0.0\% while improving aggregate utility from 46.7\% to 63.3\%. On a 200-case AgentDyn Dailylife benchmark, it reduces confirmed compromise from 73.5\% to 0.0\% while preserving near-baseline utility (44.5\% to 43.5\%). Breadth checks across ASB, InjecAgent, BIPIA, AgentHarm, and MCPTox replays suggest that the configured policies block the benchmark-specified policy-visible attacker flows; in ASB's direct-prompt-injection harness, attack success is 0/1{,}200. These results are preliminary and scoped to the modeled policy-visible agent behaviors and evaluated benchmarks.