管控智能体AI系统的执行风险:一种基于轨迹的红队测试框架
Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming
查看机构详情
- Department of Information Systems, Business Statistics and Operations Management (ISOM), Hong Kong University of Science and Technology(香港科技大学信息系统、商务统计与运营管理系(ISOM))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出基于轨迹的红队测试框架TrajRed和运行时管控层TrajGuard,在AgentDojo实验中,TrajRed识别的漏洞更强,TrajGuard可将攻击成功率降至近零且保留任务效用。
中文摘要 AI 辅助
AI智能体正越来越多地嵌入组织工作流程,在这些流程中,它们会与外部信息源交互并调用数字工具来执行操作任务。随着组织采用此类系统,一项关键挑战是识别并缓解由恶意或不可信外部信息引发的风险,这类信息可能会引导智能体采取意外行动。现有的红队测试方法大多依赖固定攻击模板或最终攻击结果,对攻击如何通过多步骤推理和工具使用展开的可见性有限。我们认为,智能体执行风险应被理解为一种轨迹层面的现象。基于这一视角,我们提出了TrajRed,一种基于轨迹的红队测试框架,该框架利用执行轨迹来发现智能体AI系统的漏洞。我们还开发了TrajGuard,这是一个运行时管控层,它利用红队测试期间发现的高风险轨迹来监控和干预正在进行的工作流程。在AgentDojo上针对四个组织任务套件开展的实验表明,TrajRed识别出的漏洞比固定模板和自动红队基线强得多。基于这些漏洞发现,TrajGuard将所有评估攻击方法的攻击成功率降至接近零,同时保留了良性任务效用。综合来看,这些结果证明,执行轨迹为智能体AI系统的红队测试和风险控制提供了实用基础。本研究强调了在组织AI部署中管控智能体执行的重要性。
英文摘要
AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital tools to perform operational tasks. As organizations adopt such systems, a critical challenge is identifying and mitigating risks arising from malicious or untrusted external information that can steer agents toward unintended actions. Existing red-teaming approaches largely rely on fixed attack templates or final attack outcomes, providing limited visibility into how attacks unfold through multi-step reasoning and tool use. We argue that agent execution risk should be understood as a trajectory-level phenomenon. Building on this perspective, we propose TrajRed, a trajectory-guided red-teaming framework that uses execution trajectories to uncover vulnerabilities in agentic AI systems. We further develop TrajGuard, a runtime governance layer that uses high-risk trajectories discovered during red teaming to monitor and intervene in ongoing workflows. Experiments on AgentDojo across four organizational task suites show that TrajRed identifies substantially stronger vulnerabilities than fixed-template and automatic red-team baselines. Building on these vulnerability findings, TrajGuard reduces attack success across all evaluated attack methods to near zero while preserving benign task utility. Together, the results demonstrate that execution trajectories provide a practical foundation for both red teaming and risk control in agentic AI systems. This work highlights the importance of governing agent execution in organizational AI deployments.