AI 中文总结
本研究通过NeuroSploit框架在GOAD环境上编排多模型进行AD渗透测试,实现96-100%技术覆盖率,显著优于直接调用LLM,验证了结构化编排的有效性。
AI 中文摘要
活动目录(AD)仍是企业环境中占主导地位的身份与访问管理基础设施,其沦陷代表着内网渗透测试中影响最严重的后果。近期研究表明,大型语言模型(LLM)能够自主对AD进行假设入侵式渗透测试,但这些研究采用独立智能体,缺乏结构化护栏、确定性验证和多阶段链式编排。我们对NeuroSploit v4.2.0(一个基于Rust的开源自主渗透测试编排框架)在Game of Active Directory(GOAD)上进行了基准评估。GOAD是由Orange Cyberdefense维护的一个刻意设计为易受攻击的多林AD实验环境,包含五台虚拟机、两个林和三个域。该框架编排了22个AD专用智能体和7个多阶段攻击链剧本,覆盖完整的AD攻击链:枚举、Kerberoasting、AS-REP roasting、NTLM中继与强制认证、Kerberos委派滥用、AD CS利用(ESC1-ESC8)、MSSQL链接服务器横向移动、DCSync、跨林信任滥用以及持久化检测。我们在该框架内对九种前沿LLM(Claude Opus 4.6/4.7/4.8、GPT-6 Astra、GPT-5.6 Sol、Grok 4.6、Qwen 3.8、GLM 5.3、Kimi k3)进行了基准测试,并与直接调用方式在14个技术类别和7条攻击链上进行了对比。该框架实现了96-100%的技术覆盖率,精确率为90-97%,而直接调用仅覆盖21-54%,且误报数量高出3.2倍。使用Opus 4.8实现三个域完全沦陷的时间为134分钟,期间护栏激活阻止了触发锁定的密码喷射、未经授权的DCSync转储和超范围侦察。结果表明,采用领域专用智能体、POMDP信念跟踪和跨模型投票的结构化框架编排在AD渗透测试中显著优于非结构化的LLM使用方式。
英文摘要
Active Directory (AD) remains the predominant identity and access management infrastructure in enterprise environments, and its compromise represents the highest-impact outcome in internal penetration tests. Recent work has shown that large language models (LLMs) can autonomously conduct assumed-breach penetration testing against AD, but these studies employ standalone agents lacking structured guardrails, deterministic validation, and multi-stage chain orchestration. We present a benchmark evaluation of NeuroSploit v4.2.0, an open-source Rust-based autonomous pentest harness, against the Game of Active Directory (GOAD), a deliberately vulnerable multi-forest AD lab maintained by Orange Cyberdefense comprising five virtual machines, two forests, and three domains. The harness orchestrates 22 AD-specific agents and 7 multi-stage attack-chain playbooks covering the full AD kill chain: enumeration, Kerberoasting, AS-REP roasting, NTLM relay and coercion, Kerberos delegation abuse, AD CS exploitation (ESC1-ESC8), MSSQL linked-server pivoting, DCSync, cross-forest trust abuse, and persistence detection. We benchmark nine frontier LLMs (Claude Opus 4.6/4.7/4.8, GPT-6 Astra, GPT-5.6 Sol, Grok 4.6, Qwen 3.8, GLM 5.3, Kimi k3) within the harness, comparing against direct invocation across 14 technique categories and 7 chains. The harness achieves 96-100% technique coverage with 90-97% precision, while direct invocation covers only 21-54% and produces 3.2x more false positives. Time to full three-domain compromise with Opus 4.8 was 134 minutes with guardrail activations preventing lockout-triggering sprays, unauthorized DCSync dumps, and out-of-scope reconnaissance. Results demonstrate that structured harness orchestration with domain-specialized agents, POMDP belief tracking, and cross-model voting substantially outperforms unstructured LLM usage for AD penetration testing.
Comments20 pages, 15 figures, 7 tables, 4 listings, 23 references