发表机构
Tsinghua University; CRRC Corporation Limited(清华大学; 中国中车股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LLM智能体攻击链重构基准的不足,提出DiagChain基准与ECRAG方法,经6种LLM评估发现模型在证据整合与排序环节存在瓶颈,为网络安全智能体改进提供了诊断依据。
AI 中文摘要
大语言模型(LLM)智能体为攻击链重构提供了一种有前景的方法,可通过检索和解释异构遥测数据来推断攻击者的有序操作。然而,现有基准主要评估最终输出或聚合准确率,对错误如何在中间推理阶段产生和传播的洞察有限。我们提出DiagChain,这是一个用于基于证据的攻击链重构的诊断基准,支持对LLM智能体进行分阶段评估。DiagChain包含MAIN-69,这是一套69个场景,涵盖多个操作系统、证据噪声水平和链长度。它还引入了以证据为中心的检索增强生成(ECRAG),将证据检索与重构链的演化结构化表示相结合。我们引入了五个互补指标,以评估重构过程的不同阶段并支持系统性故障诊断。基于对6个LLM的评估,DiagChain显示,即使是最强的配置在MAIN-69的849个参考步骤中也仅成功39.6%。我们的分析进一步表明,较小的模型在将检索到的证据纳入输出的更基础任务上存在困难,而较大的模型可以推进到后续步骤,此时正确排序该证据成为主要瓶颈。这些结果验证了超越端到端准确率进行诊断评估的重要性,并为改进基于证据的网络安全智能体提供了可操作的见解。
英文摘要
Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to infer ordered attacker actions. However, existing benchmarks mainly evaluate final outputs or aggregate accuracy, providing limited insight into how errors arise and propagate across intermediate reasoning stages. We present DiagChain, a diagnostic benchmark for evidence-grounded attack chain reconstruction that enables stage-wise evaluation of LLM agents. DiagChain includes MAIN-69, a suite of 69 scenarios spanning multiple operating systems, evidence noise levels, and chain lengths. It further introduces Evidence-Centric Retrieval-Augmented Generation (ECRAG), which couples evidence retrieval with an evolving structured representation of the reconstructed chain. Five complementary metrics are introduced to assess distinct stages of the reconstruction process and support systematic failure diagnosis. Based on evaluations using 6 LLMs, DiagChain reveals that even the strongest configuration succeeds on only 39.6% of the 849 reference steps in MAIN-69. Our analysis further shows that smaller models struggle with the more basic task of incorporating retrieved evidence into their outputs, whereas larger models can proceed to later steps, where correctly ordering that evidence becomes the main bottleneck. These results validate the importance of diagnostic evaluation beyond end-to-end accuracy and provide actionable insights for improving evidence-grounded cybersecurity agents.