arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16694cs.CR

迈向安全的AI驱动渗透测试代理:安全威胁、防护机制与架构视角

Toward Secure AI-Powered Penetration Testing Agents: Security Threats, Guardrails, and Architectural Perspectives

发表机构印度国家技术学院卡利卡特分校
查看机构详情
  • National Institute of Technology Calicut(印度国家技术学院卡利卡特分校)

机构由 AI 辅助整理,请以论文原文为准。

Rahul Dev T Y, Hiran V Nath

首次发表
浏览论文内容

中文总结 AI 辅助

该研究系统分析AI渗透测试代理的安全威胁,提出涵盖LLM生命周期、架构及行为的威胁分类法,并指出现有防护不足,倡导开发上下文与架构感知的专门防护机制。

中文摘要 AI 辅助

基于大语言模型(LLM)的自主代理正以动态、多步骤的进攻性安全工作流变革渗透测试领域,这些工作流仅需最少的人类监督。这些代理利用复杂的推理能力和外部安全工具,独立执行侦察、识别漏洞、制定利用计划并进行后渗透操作。然而,与传统的基于聊天的LLM系统相比,具备持久记忆、在现实世界中采取行动以及进行长时程推理的能力引发了性质不同的安全担忧。现有的针对对话式AI的防护机制可能不足以相应地保护自主AI渗透测试代理。为解决这些问题,我们对自主AI渗透测试代理进行了全面的安全分析。我们系统地分析了代表性的代理架构,刻画了其信任边界和攻击面,并提出了一种与生命周期对齐的威胁分类法,该分类法涵盖LLM生命周期攻击、代理架构攻击和跨领域行为攻击。我们分析了现有防护机制的局限性,识别了关键研究空白,并讨论了开发专门的、上下文感知的和架构感知的防护机制以保护下一代AI驱动的进攻性安全系统的未来研究方向。

英文摘要

LLM-powered autonomous agents are transforming the penetration testing space with dynamic, multi-step offensive security workflows that require minimal supervision by humans. These agents leverage sophisticated reasoning abilities and external security tools to independently carry out reconnaissance, identify vulnerabilities, devise exploitation plans, and perform post-exploitation operations. But the ability to have persistent memory, to take actions in the real world, and to do long-horizon reasoning raises qualitatively different security concerns than traditional chat-based LLM systems. Existing guardrail mechanisms for conversational AI may not be sufficient to secure autonomous AI pentesting agents accordingly. To address these issues, we carry out a comprehensive security analysis on autonomous AI-penetration testing agents. We systematically analyse representative agent architectures, characterise their trust boundaries and attack surfaces and propose a threat taxonomy that is aligned with the lifecycle and covers LLM lifecycle attacks, agent-architecture attacks and cross-cutting behavioural attacks. We analyse the limitations of existing guardrail mechanisms, identify key research gaps, and discuss future research directions for developing specialised, context-aware, and architecture-aware guardrails to secure next-generation AI-driven offensive security systems.

↑