发表机构
Tongyi Lab, Alibaba Group; Alibaba Cloud Computing, Alibaba Group; The Hong Kong University of Science and Technology(阿里巴巴集团通义实验室; 阿里巴巴集团阿里云计算; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人员推出首个入侵后事件响应AI智能体基准测试集SecRespond,评估23种前沿LLM在10个靶场的表现,发现现有智能体难以及时发现隐蔽入侵并制定全面修复计划,存在根本瓶颈。
AI 中文摘要
大型语言模型(LLM)智能体正越来越多地被应用于具备主机制品和命令行界面(CLI)的真实安全运营场景,因此对其安全能力进行全面评估至关重要。然而,现有的网络安全基准测试集均聚焦于入侵前场景,即智能体在攻击发生前被置于干净、理想化的环境中,这使得入侵后场景的研究被忽视。为填补这一空白,我们推出了SecRespond——首个用于评估LLM智能体入侵后事件响应工作流的基准测试集。给定被入侵主机的取证磁盘快照,以及主机安全产品报告的警报、漏洞扫描结果和基线检查结果,智能体需生成关于入侵、基线风险和漏洞风险的取证报告,并制定修复计划。我们在10个网络靶场中实例化该任务,每个靶场由不同的被入侵云主机构建,涵盖4种入口点类型、21种ATT&CK技术和5种操作系统。我们在OpenCode智能体框架上评估了23种前沿LLM。实验结果表明,尽管当前智能体能可靠地发现警报暴露的问题,但它们难以主动调查磁盘中的隐蔽入侵,也难以生成全面、经核实的修复计划,没有任何模型在单个靶场上实现完整的检测与修复。这揭示了构建用于真实事件响应的智能体存在的根本瓶颈,该基准测试集可通过指定URL公开获取。
英文摘要
Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise settings where agents are placed in a clean and idealized environment before an attack occurs. This leaves the post-compromise setting underexplored. To address this gap, we introduce SecRespond, the first benchmark for evaluating LLM agents on the post-compromise incident-response workflow. Given a forensic disk snapshot of a compromised host together with the alerts, vulnerability scans, and baseline checks reported by a host security product, agents are required to produce forensic reports on intrusions, baseline risks, and vulnerability risks, together with a remediation plan. We instantiate this task across 10 cyber ranges, each constructed from a distinct compromised cloud host, spanning 4 entry-point types, 21 ATT&CK techniques, and 5 operating systems. We evaluate 23 frontier LLMs on the OpenCode agent harness. Experimental results show that although current agents can reliably uncover the problems exposed by alerts, they struggle to proactively investigate the disk for silent intrusions and to produce comprehensive, verified remediation plans, with no model achieving complete detection and remediation on any single range. This reveals a fundamental bottleneck in building agents for real-world incident response. The benchmark is publicly available at https://github.com/Alibaba-NLP/qqr/tree/main/data/secrespond.