OpenAgentFlow:为异构AI智能体集群提供系统级安全边界
OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
浏览论文内容
中文总结 AI 辅助
本文提出OpenAgentFlow架构,在Android上实现异构AI智能体集群的系统级安全边界,攻击拦截率95.3%,可支持动态策略更新,提升安全管控能力。
中文摘要 AI 辅助
由大语言模型驱动的AI智能体正从孤立助手演变为异构系统,其中多个智能体、规划器、控制器和执行后端在同一用户或企业环境中运行。在此类场景中,安全成为系统级的行动治理问题:需在智能体生成的具体行动修改共享状态前,决定是否应提交该行动。现有安全措施覆盖提示词、工具调用、图形界面(GUI)行动及智能体本地行为,但常存在执行碎片化、掩盖多步骤行动流中出现的风险、对可审计性和策略演进支持有限等问题。本文提出OpenAgentFlow,一种在行动提交边界执行安全控制的控制平面/行动平面架构。它将待执行的GUI行动、API调用、工具调用及大语言模型(LLM)生成的调用归一化为统一的AgentEvent流,将每个事件路由至共享的执行前策略执行点(Policy Enforcement Point),并在控制平面维护溯源信息、会话状态、审计记录及可更新策略。这一架构创建了可共享治理的行动流,且无需修改智能体、提示词、模型或执行路径即可让新规则生效。我们在Android上实例化OpenAgentFlow,在含300个行动事件的基准测试中,其准确率达94.0%,攻击拦截率达95.3%;在含30个动态策略的测试套件中,安装新规则后27个案例符合预期行为;在100个案例的Android模拟器套件中追踪的98个案例里,其在GUI、API及LLM规划案例中的原始准确率达90.8%,经追踪调整的通过率达92.9%。这些结果表明,OpenAgentFlow为异构AI智能体集群提供了实用的共享执行边界。
英文摘要
AI agents powered by large language models are evolving from isolated assistants into heterogeneous systems in which multiple agents, planners, tools, and execution backends operate over shared environments. In such settings, safety becomes a system-level action-governance problem: deciding whether a pending action should be committed given policy-relevant state accumulated across a session. Existing safeguards operate at fragmented boundaries, making it difficult to enforce shared policies over composed action flows across heterogeneous execution paths. We present OpenAgentFlow, a control-plane/action-plane architecture that establishes the action-commit boundary as a shared enforcement interface. GUI, API, tool, and LLM-generated actions are normalized into a common AgentEvent stream and mediated by a shared pre-execution Policy Enforcement Point, while provenance, session state, audit evidence, and updatable policies are maintained outside individual agents. This provides a common governance layer across incompatible executors and allows new policies to take effect without modifying agents, prompts, models, or execution paths. We evaluate OpenAgentFlow through complementary system evaluations spanning controlled action-flow tests, a public external benchmark, policy updates, and real Android execution. On a 300-case controlled suite, OpenAgentFlow achieves 94.00% accuracy and a 95.35% attack-block rate. On the complete 1,220-case AgentDojo-Traj split of TS-Bench, it achieves 97.62% accuracy, 96.59% unsafe-action recall, and a 1.96% safe false-intervention rate. New control-plane rules take effect without modifying protected agents, and the same enforcement path operates across live GUI, API/tool, and LLM-planned Android execution. These results show that a shared action-commit boundary provides a practical basis for system-wide governance across heterogeneous agent execution paths.
发表机构
- Southern University of Science and Technology(南方科技大学)
- City University of Hong Kong(香港城市大学)
- Lingnan University(岭南大学)
机构由 AI 辅助整理,请以论文原文为准。