arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09647cs.AI

智能体AI的黑盒红队测试:基于分类学的自动化风险发现框架

Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery

Divyanshu Kumar, Nitin Aravind Birur, Tanay Baswa, Sahil Agarwal, Prashanth Harshangi

首次发表
浏览论文内容

中文总结 AI 辅助

提出基于七域分类学的黑盒红队框架SAGE-RT,自动生成对抗场景并评估智能体系统,揭示治理、隐私及行为漏洞的高风险率。

中文摘要 AI 辅助

智能体系统正迅速投入生产环境,它们读取不可信输入、调用具有真实权限的工具并自主行动,将安全边界扩展到了仅限聊天的模型之外。然而,标准评估仍停留在单轮交互层面,无法捕捉多步骤的智能体漏洞。我们提出了一个系统的黑盒框架,用于风险感知的智能体评估,该框架仅需基本的系统描述。我们的方法引入了:(1)一个七域分类学,将可观察行为映射到风险类别;(2)完全自动化的SAGE-RT红队测试,为每个域生成120个对抗性场景;(3)使用LLM评判员进行人工验证的评估。在两种智能体架构(CrewAI和AutoGen)上使用四个基础模型进行的实证验证揭示了令人担忧的模式:平均治理风险为56.25%,多智能体配置中的隐私风险为65%,智能体行为漏洞高达85%。我们的黑盒方法无需特权访问即可有效识别关键架构漏洞,为更安全的智能体部署提供了一条可扩展的路径。

英文摘要

Agentic systems are rapidly moving to production, where they read untrusted inputs, call tools with real permissions, and act autonomously, expanding the security surface beyond chat-only models. Yet standard evaluations remain single-turn and fail to capture multi-step agent vulnerabilities. We present a systematic black-box framework for risk-aware agent evaluation requiring only basic system descriptions. Our approach introduces: (1) a seven-domain taxonomy mapping observable behaviors to risk categories, (2) fully automated SAGE-RT red teaming producing 120 adversarial scenarios per domain, and (3) human-validated evaluation using LLM judges. Empirical validation across two agent architectures (CrewAI and AutoGen) with four base models reveals alarming patterns: 56.25\% average governance risk, 65\% privacy risk in multi-agent configurations, and agent behavior vulnerabilities reaching 85\%. Our black-box approach effectively identifies critical architectural vulnerabilities without privileged access, providing a scalable path toward safer agent deployments.

发表机构

  • Enkrypt AI

机构由 AI 辅助整理,请以论文原文为准。

↑