arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18120cs.CRcs.AIcs.NI

PentestChain:一种成本感知的、基于MCP编排的自动化渗透测试框架,使用免费层LLM

PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs

Rushabh Vipulkumar Patel, Dipo Dunsin, Mohammed Almaiah, Mohamed Chahine Ghanem

首次发表
浏览论文内容

中文总结 AI 辅助

PentestChain提出一种成本感知的自动化渗透测试框架,通过确定性图谱与免费层LLM级联实现零API成本,并分析MCP攻击面,提供可重现的评估协议。

中文摘要 AI 辅助

AI驱动的渗透测试已通过GPT-4等高级前沿模型得到展示,但每次参与(engagement)的令牌成本使得持续、自动化的测试对于最需要它的小型组织来说负担不起。本文提出了PentestChain,一个十阶段的自动化渗透测试框架,它将一个精心策划的确定性漏洞利用图谱与一个成本感知的AI级联相结合——首先是本地Ollama模型(qwen2.5-7b),然后是免费层的OpenRouter和Cerebras,并有一个基于规则的备选方案,该方案始终产生输出——并通过一个具有十一个工具的模型上下文协议(MCP)服务器暴露整个流水线。我们做出了三项贡献。首先,我们将每次参与的美元成本视为一个可测量的、一流的评估指标,并表明一个7B参数的本地模型,通过确定性主干保持在关键路径之外,能够在零实测付费API成本的情况下维持端到端操作。其次,我们分析了MCP暴露的攻击引擎引入的攻击面,将四位置威胁模型建立在2025年MCP事件记录上(mcp-remote中的CVE-2025-6514远程代码执行漏洞、postmark-mcp供应链后门以及工具投毒-拉地毯-跳线类),并贡献了四项缓解措施。第三,我们指定了一个可重现的、容器化的评估协议,与顶级场所现在期望的标准化测试平台保持一致——AutoPenBench、Cybench子集和PentestGPT 182子任务基准——具有多试验统计(每个配置超过10次试验、pass-at-k、非参数显著性检验和效应量)以及直接在同一测试平台上重现PentestGPT和PentestAgent基线,而不是引用其已发表的数据。在迄今为止测量的遗留目标上,该框架检测了26项服务,丰富了34个CVE,产生了

英文摘要

AI-driven penetration testing has been demonstrated with premium frontier models such as GPT-4, but the per-engagement token cost makes continuous, automated testing unaffordable for the smaller organisations that need it most. This paper presents PentestChain, a ten-phase automated penetration testing framework that couples a curated, deterministic exploit map with a cost-aware AI cascade-a local Ollama model (qwen2.5-7b) first, then free-tier OpenRouter and Cerebras, with a rule-based fallback that always produces output-and exposes the full pipeline through a Model Context Protocol (MCP) server with eleven tools. We make three contributions. First, we treat US-dollar cost per engagement as a measured, first-class evaluation metric and show that a 7B-parameter local model, kept off the critical path by a deterministic backbone, sustains end-to-end operation at zero measured paid-API cost. Second, we analyse the attack surface that an MCP-exposed offensive engine introduces, grounding a four-position threat model in the 2025 MCP incident record (the CVE-2025-6514 remote-code-execution flaw in mcp-remote, the postmark-mcp supply-chain backdoor, and the tool-poisoning-rug-pull-line-jumping class), and contribute four mitigations. Third, we specify a reproducible, containerised evalua-tion protocol aligned with the standardised testbeds now expected at top-tier venues-AutoPenBench, a Cybench subset, and the PentestGPT 182-sub-task benchmark-with multi-trial statistics (more than 10 trials per configuration, pass-at-k, non-parametric significance tests and effect sizes) and direct, same testbed reproduction of the PentestGPT and PentestAgent baselines rather than citation of their published numbers. On the legacy targets measured to date, the framework detected 26 services, enriched 34 CVEs, produced

发表机构

  • University of Wales Trinity Saint David(威尔士三一圣大卫大学)
  • The University of Jordan(约旦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑