arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07344cs.CRcs.AI

保持在攻击路径上:面向长时程自动化渗透测试的结构化状态

Staying on the Attack Path: Structured State for Long-Horizon Automated Penetration Testing

  • Tianjin University(天津大学)
  • Changchun University of Science and Technology(长春理工大学)
  • Jilin Province Key Laboratory of Network and Information Security(吉林省网络与信息安全重点实验室)
  • Shihezi University(石河子大学)

机构由 AI 辅助整理,请以论文原文为准。

Weizhe Wang, Yitong Zhang, Yao Zhang, Xiaoqiang Di, Zhigang Li, Bin Wu, Guangquan Xu

中文总结 AI 辅助

本文提出Intentest,一种基于意图图引导的自动化渗透测试智能体,通过外部化长时程状态至事实-意图DAG,在CTF基准上实现88.2%成功率,显著优于基线。

中文摘要 AI 辅助

基于大型语言模型(LLM)的智能体正越来越多地应用于网络安全任务,如漏洞发现和自动化渗透测试。然而,在长时程安全任务中,此类智能体仍受限于上下文遗忘和意图漂移:早期关键事实和因果推理链在长时间交互中丢失,智能体陷入无目的的重复探索。本文提出Intentest,一种意图图引导的自动化渗透测试智能体,将长时程状态从LLM的上下文窗口外部化到持久的事实-意图有向无环图(DAG)上,从而大幅减少无效转换。我们在Web应用自动化渗透测试上评估Intentest,这是网络安全中一项具有代表性的长尾任务。在DAG中,已验证的网络状态存储为不可变事实节点,探索方向被约束为由前驱事实限定的意图边。系统采用三层架构,其中事实-意图映射层维护全局状态,任务调度与分配层通过两阶段降级恢复和多维自适应负载均衡确保执行稳定性,意图检索与预测层通过自顶向下的五阶段过滤算法提供战术先验。在一个涵盖三种难度级别、超过十种漏洞类型的真实CTF挑战基准上,Intentest实现了88.2%的总体成功率和75.0%的困难任务成功率,相比基线分别提升约44和50个百分点。消融实验进一步表明,意图检索与预测在不改变可解任务集的情况下,将成功的中等和困难任务的平均轮数分别减少约33%和48%。

英文摘要

Large language model (LLM) based agents are increasingly applied to cybersecurity tasks such as vulnerability discovery and automated penetration testing. On long-horizon security tasks, however, such agents remain limited by context forgetting and intent drift: early critical facts and causal reasoning chains are lost over extended interactions, and the agent falls into aimless, repetitive exploration. This paper proposes Intentest, an intent-graph-guided automated penetration testing agent that externalizes long-horizon state from the LLM's context window onto a persistent fact-intent directed acyclic graph (DAG), thereby substantially reducing invalid transitions. We evaluate Intentest on automated penetration testing of web applications, a representative long-tail task in cybersecurity. In the DAG, verified network states are stored as immutable fact nodes, and exploration directions are constrained as intent edges bounded by predecessor facts. The system adopts a three-layer architecture, in which the fact-intent mapping layer maintains the global state, the task scheduling and allocation layer ensures execution stability through two-phase degradation recovery and multi-dimensional adaptive load balancing, and the intent retrieval and prediction layer provides tactical priors through a top-down five-stage filtering algorithm. On a benchmark of real CTF challenges covering more than ten vulnerability types across three difficulty levels, Intentest achieves an overall success rate of 88.2% and a success rate of 75.0% on hard tasks, improving over the baseline by approximately 44 and 50 percentage points. Ablation experiments further show that the intent retrieval and prediction reduce the average number of rounds on successful medium and hard tasks by about 33% and 48%, respectively, without changing the set of solvable tasks.

↑