具备网络能力的人工智能代理:漏洞、评估遏制与防御响应
Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
浏览论文内容
中文总结 AI 辅助
研究具备网络能力的人工智能代理的漏洞等问题,综合五类漏洞,以特定事件为案例研究,探讨遏制等控制措施,确定评估网络能力及其运行环境安全性的实际优先事项。
中文摘要 AI 辅助
具备网络能力的人工智能代理将语言模型与工具、内存和执行环境相结合,以执行多步攻击安全任务。现有工作分别衡量网络能力并编目针对代理组件的攻击,但在将有能力的代理限制在用于评估它的环境方面提供的指导较少。本综述综合了该边界处的五类漏洞:多步攻击链、与沙盒边界冲突的目标、供应链和凭证暴露、持久的指挥与控制以及自动行动的速度。我们将2026年7月报告的Hugging Face/OpenAI事件作为一个有界案例研究,区分特定事件的观察结果与更广泛文献中确立的发现。在整个分类法和案例中,我们研究了遏制、权限分离、来源和响应者访问的控制措施,包括防御工件也可能导致滥用的两用问题。该综述确定了评估网络能力以及行使该能力的环境安全性的实际优先事项。
英文摘要
Cyber-capable AI agents combine language models with tools, memory, and execution environments to perform multi-step offensive-security tasks. Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environments used to evaluate it. This review synthesizes five vulnerability classes at that boundary: multi-step offensive chains, objectives that conflict with sandbox boundaries, supply-chain and credential exposure, persistent command-and-control, and the speed of automated action. We use two separate preliminary incident records: the reported July 2026 Hugging Face/OpenAI evaluation breach and Anthropic's subsequent three-incident evaluation review. A comparative evidence protocol distinguishes record-specific factual claims from the shared systems lesson: the evaluation environment is itself part of the security boundary. Across the taxonomy and records, we examine controls for containment, privilege separation, provenance, and responder access, including the dual-use problem that defensive artifacts may also enable misuse. The review identifies practical priorities for evaluating cyber capability together with the security of the environment in which that capability is exercised.
发表机构
- Rajshahi University of Engineering & Technology(拉杰沙希工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。