arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32424cs.CRcs.AI

CyberClear:面向APT攻击链溯源的大语言模型智能体系统基准

CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance

Qi Chen, Fushuo Huo, Hangli Shen, Jingcai Guo, Shuhao Li, Guang Cheng

首次发表
浏览论文内容

中文总结 AI 辅助

提出CyberClear基准,评估大语言模型智能体从长上下文安全日志中重建APT攻击链的能力,并设计CyberProvenance框架,通过证据积累与执行验证提升多智能体系统溯源性能。

中文摘要 AI 辅助

大语言模型智能体在网络安全任务中已展现出令人瞩目的能力,然而,它们从复杂安全日志中重建完整高级持续性威胁(APT)攻击活动的能力在很大程度上仍未得到探索。现有的面向智能体的网络安全基准主要聚焦于漏洞发现、利用和安全分析任务,导致在真实安全日志下对攻击链溯源的评估研究不足。为填补这一空白,我们提出了CyberClear,一个用于评估大语言模型智能体及先进智能体系统在长上下文安全日志中进行APT攻击链溯源能力的基准。CyberClear涵盖单步攻击和多阶段攻击链,要求智能体识别攻击证据、推断攻击进程,并生成包含实体、因果关系、MITRE ATT&CK技术及取证证据的溯源图。为支持全面评估,我们开发了一种针对APT攻击链溯源量身定制的评估方法。与聚焦于表面匹配的传统文本相似度指标不同,我们的评估考察重建图是否在单步行为正确性、多步行为识别、时间与因果一致性、实体与关系保真度以及整体攻击叙事一致性等方面保留攻击链的语义。由最先进大语言模型驱动的先进多智能体系统在CyberClear上仍表现挣扎,这促使我们提出CyberProvenance,一个专为多智能体设计的智能体网络防护框架,通过证据积累、基于执行的验证和反馈引导的细化机制增强大语言模型智能体,以实现可靠的攻击链溯源。在CyberClear上的广泛评估证明了CyberProvenance在改进证据推理、基于执行的验证以及完整APT攻击链重建方面的有效性。

英文摘要

Large language model agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulnerability discovery, exploitation, and security analysis tasks, leaving the evaluation of attack chain provenance under realistic security logs insufficiently studied. To address this gap, we introduce CyberClear, a benchmark for evaluating LLM agents and advanced agent systems on APT attack chain provenance from long-context security logs. CyberClear covers both single-step attacks and multi-stage attack chains, requiring agents to identify attack evidence, infer attack progression, and generate provenance graphs containing entities, causal relationships, MITRE ATT&CK techniques, and forensic evidence. To enable comprehensive evaluation, we develop an evaluation method tailored to APT attack chain provenance. Unlike conventional text similarity metrics that focus on surface-level matching, our evaluation examines whether reconstructed graphs preserve the semantics of attack chains across single-step behavior correctness, multi-step behavior identification, temporal and causal consistency, entity and relationship fidelity, and overall attack narrative consistency. Advanced multi-agent systems powered by state-of-the-art LLMs still struggle on CyberClear, motivating us to propose CyberProvenance, an agent cyber harness designed for multi-agents that augments LLM agents with evidence accumulation, execution-based validation, and feedback-guided refinement mechanisms for reliable attack-chain provenance. Extensive evaluations on CyberClear demonstrate the effectiveness of CyberProvenance in improving evidence reasoning, execution-grounded validation, and complete APT attack chain reconstruction.

发表机构

  • Southeast University(东南大学)
  • The Hong Kong Polytechnic University(香港理工大学)
  • Zhongguancun Laboratory(中关村实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑