arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动化攻击图构建用于智能体渗透测试:迈向神经-符号漏洞 hunting

Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting

Oliver Stevanovic, Jasmin Wachter

arXiv 2609.15523首次发表:更新:

发表机构

University of Udine; University of Klagenfurt(乌迪内大学; 克拉根福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出半自动管道,将Trivy、Semgrep和Nmap扫描结果解析为MulVAL谓词,借助LLM构建Datalog规则,通过符号推理生成攻击图,在54个Web夺旗任务中实现平均53.7%的真实漏洞覆盖率,验证了其在智能体渗透测试中的可行性。

AI 中文摘要

基于扫描器输出的逻辑攻击图提供了显式且可审计的攻击路径推理,这是基于LLM的智能体所缺乏的。然而,将诸如MulVAL之类的符号框架集成到当代安全工作流或智能体管道中,需要将扫描器证据转换为初始事实,并创建领域特定规则。我们提出了一个半自动管道来解决这一互操作性问题,并在一个Web安全案例研究中展示了其可行性。我们的管道将Trivy、Semgrep和Nmap的发现解析为MulVAL谓词,并使用LLM辅助过程构建领域特定的Datalog规则,将扫描器可检测的证据与攻击技术联系起来。然后,MulVAL/XSB执行符号推理以生成结构化的攻击路径。我们在一个智能体管道(混合推理器)中,对来自CyBench的54个Web夺旗任务评估了攻击图构建基础设施;我们不评估下游智能体的性能。每个任务至少产生一个达到目标的图,我们实现了平均真实漏洞覆盖率为53.7%,其中51.9%的任务实现了完全覆盖;平均噪声路径率为83.9%。端到端时间中位数为24.9秒(MulVAL推理:2.7秒),该管道对于智能体工作流是可行且运行时实用的,但谓词覆盖、规则覆盖和路径精度仍是限制因素。下一步包括语义规则验证和用于图引导渗透测试的智能体级比较。

英文摘要

Logic attack graphs grounded in scanner output provide explicit and auditable attack path reasoning LLM-based agents lack. Integrating symbolic frameworks such as MulVAL to contemporary security workflows or agentic pipelines, however, requires translating scanner evidence to initial facts, and creating domain-specific rules. We present a semi-automated pipeline that addresses this interoperability problem and depict its feasibility in a web-security case study. Our pipeline parses findings from Trivy, Semgrep, and Nmap into MulVAL predicates and uses an LLM-assisted process to construct domain-specific Datalog rules linking scanner-detectable evidence to attack techniques. MulVAL/XSB then performs symbolic inference to generate structured attack paths. We evaluate the attack-graph construction infrastructure on 54 web Capture-the-Flag tasks from CyBench within an agentic pipeline (Hybrid Reasoner); we do not evaluate the performance of the downstream agent. Every task produced at least one goal-reaching graph, and we achieve mean ground-truth vulnerability coverage of 53.7%, with 51.9% achieving full coverage; mean noise-path rate was 83.9%. With median end-to-end time of 24.9 s (MulVAL reasoning: 2.7 s) the pipeline is feasible and runtime-practical for agentic workflows, but predicate coverage, rule coverage, and path precision remain limiting factors. Next steps include semantic rule validation and agent-level comparison for graph-guided pentesting.

CommentsCite as: Stevanovic, O., & Wachter, J. (2026). Automating attack graph construction for agentic pentesting: Towards neuro-symbolic vulnerability hunting. In D. Hitaj et al. (Eds.), ESORICS 2026 workshops. Springer Nature Switzerland AG

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑