arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MazeRunner:面向大语言模型驱动的黑盒自动化渗透测试的非线性任务与线索编排

MazeRunner: Nonlinear Task and Clue Orchestration for LLM-driven Black-Box Automated Penetration Testing

Zhenyuan Li, Yi Jiang, Junjie Cheng, Yaokun Li, Jing Qiu, Shouling Ji

arXiv 2608.14216首次发表:更新:

AI 中文总结

针对现有LLM驱动渗透测试智能体在非线性黑盒场景中的缺陷,提出三智能体框架的MazeRunner,在HTB目标上完成更多子任务、获取更高权限,探索攻击分支更高效。

AI 中文摘要

渗透测试至关重要但资源密集。尽管大语言模型(LLM)展现出自动化安全审计的潜力,现有智能体主要在简化的线性场景中执行端到端工作流。现实世界的黑盒测试本质上是非线性的:攻击图初始未知,必须从环境反馈中逐步推断;观测结果可能揭示多个攻击分支,失败原因往往模糊不清,关键线索可能跨越很长的行动时间范围。因此,现有智能体容易陷入深度优先探索、错误诊断失败、遗忘先前证据的困境。我们提出MazeRunner,一种基于三智能体任务与线索编排框架构建的自主渗透测试系统,它将全局编排、上下文密集型执行和面向失败的审查相分离,同时持续维护任务状态和环境证据。该设计支持行动修正、前置条件恢复、分支切换和远程线索关联。我们在10个最新发布的HTB目标上评估MazeRunner,将每个系统-目标运行限制为2000万LLM token,并防止目标特定解决方案泄露。使用Claude Sonnet 4.5时,MazeRunner完成47.7%的标注子任务,而PentestGPT-V2为36.2%,Claude Code为34.2%;它在6个目标上获得用户级或更高权限,包括2个目标的root权限,而相同模型的基线仅在2个目标上达到用户级权限,从未获得root权限。执行轨迹分析进一步表明,MazeRunner探索更多攻击分支,且更高效地获取shell。

英文摘要

Penetration testing is essential yet resource-intensive. Although large language models (LLMs) show promise for automating security auditing, existing agents mainly execute end-to-end workflows in simplified linear scenarios. Real-world black-box testing is fundamentally nonlinear: the attack graph is initially unknown and must be incrementally inferred from environmental feedback. Observations may reveal multiple attack branches, failures are often ambiguous, and critical clues may span long action horizons. Existing agents therefore tend to become trapped in depth-first exploration, misdiagnose failures, and forget prior evidence. We present MazeRunner, an autonomous penetration testing system built on a three-agent task-and-clue orchestration framework. It separates global orchestration, context-intensive execution, and failure-oriented review while persistently maintaining task states and environmental evidence. This design supports action revision, prerequisite recovery, branch switching, and long-range clue correlation. We evaluate MazeRunner on 10 recently released HTB targets, limiting each system-target run to 20 million LLM tokens and preventing target-specific solution leakage. With Claude Sonnet 4.5, MazeRunner completes 47.7% of annotated subtasks, compared with 36.2% for PentestGPT-V2 and 34.2% for Claude Code. It achieves user-level or higher access on six targets, including root access on two; each same-model baseline reaches user-level access on only two targets and never obtains root access. Execution-trace analysis further shows that MazeRunner explores more attack branches and acquires shells more efficiently.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑