发表机构
Virginia Tech; WashU(弗吉尼亚理工大学; 华盛顿大学圣路易斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Skynet将智能体执行转为工作流图,联合建模语义与结构,仅用良性数据训练,实现工作流级异常检测,在三个基准上误报率低于1%。
AI 中文摘要
智能体AI系统通过规划、工具使用和多智能体协调的长时程工作流来执行复杂任务。这些系统中的任务失败通常源于单个步骤,例如注入的提示或存在缺陷的计划,然后随着被破坏的步骤在后续众多智能体和工具调用中传播,通过下游依赖关系被放大。现有防御措施要么针对特定类别的攻击或失败,要么孤立地检查单个提示和步骤。两者都未检查工作流的全局依赖结构,并且遗漏了仅在将执行视为整体时才会出现的不一致性。我们认为,智能体AI的异常检测必须在工作流级别进行推理,因为全局执行结构暴露了局部检查无法看到的信号。我们提出了Skynet,一个原则性的工作流级异常检测框架,它将观察到的多智能体执行转换为有向工作流图,并根据学习到的良性行为对其进行评分。Skynet联合建模语义执行上下文以及智能体间委派、工具调用和数据流依赖的结构组织,并且仅在良性工作流上训练。由于训练从未见过攻击或失败,这种设计自然扩展到零日检测:任何违反良性工作流规律性的执行都会在单一决策规则下表现为流形外几何。我们在三个公开的智能体安全性和失败基准上评估了Skynet。它在保持高召回率的同时实现了低于1%的误报率,并且每个工作流和每个步骤的延迟足够低,适用于智能体AI运行时的在线监控。
英文摘要
Agentic AI systems execute complex tasks through long-horizon workflows of planning, tool use, and multi-agent coordination. Task failures in these systems often originate from a single step, such as an injected prompt or a flawed plan, and are then amplified through downstream dependencies as the corrupted step propagates across many subsequent agents and tool calls. Existing defenses either target a specific class of attacks or failures, or inspect individual prompts and steps in isolation. Both leave the global dependency structure of a workflow unexamined, and miss the inconsistencies that only emerge when the execution is viewed as a whole. We argue that anomaly detection for agentic AI must reason at the workflow level, where global execution structure exposes signals that local checks cannot see. We present Skynet, a principled workflow-level anomaly detection framework that turns observed multi-agent execution into directed workflow graphs and scores them against learned benign behavior. Skynet jointly models the semantic execution context and the structural organization of inter-agent delegation, tool invocation, and data-flow dependencies, and trains only on benign workflows. Because training never sees attacks or failures, this design naturally extends to zero-day detection: any execution that violates benign workflow regularities surfaces as off-manifold geometry under a single decision rule. We evaluate Skynet on three public agentic safety and failure benchmarks. It sustains high recall together with a sub-1% false positive rate, with per-workflow and per-step latencies low enough for online monitoring of agentic AI runtimes.
Comments10 pages, 4 figures. Accepted by ACM MobiHoc 2026