arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CyberPersistBench:评估基于LLM的网络攻击者的安装与持久化能力

CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence

Sujin Chen, Lijun Li, Xuhong Wang, Jing Shao

arXiv 2609.36573首次发表:更新:

发表机构

Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有网络安全基准仅评估初始入侵的局限,提出CyberPersistBench基准,将持久化定义为对抗性生存任务,通过六级评分和203个任务评估五个前沿LLM智能体,发现其自主持久化成功率有限(27.6%-44.8%),揭示新兴攻击风险。

AI 中文摘要

虽然基于LLM的攻击者在漏洞利用方面表现出日益增强的能力,但现有的大多数网络安全基准测试存在单阶段截断的问题,即在获得初始访问权限后便过早终止评估。在实践中,初始立足点在服务重启和主机重启等操作中断中异常脆弱。基于LLM的攻击者能否在初始入侵之外建立并维持持久的立足点,仍是网络安全评估中的一个核心盲点。为弥补这一空白,我们引入了CyberPersistBench,这是首个专门针对入侵后安装与持久化的基准测试。CyberPersistBench与前期利用解耦,将持久化定义为一种对抗性生存任务,其中智能体使用宿主原生机制在分阶段的系统中断中维持立足点。确定性检查支持六级评分方法(L1--L6),涵盖安装和持久化。该基准包含七个类别的203个核心任务,并辅以多主机和主动防御扩展。对五个前沿智能体的实证评估表明,自主持久化能力仍然有限(27.6%--44.8%),且在启用防御的任务中进一步下降(5.5%--13.3%);尽管如此,这些结果揭示了新兴的网络攻击风险,凸显了对入侵后持久化进行基准测试的必要性。因此,CyberPersistBench为入侵后安装与持久化建立了一个基础基准,划定了自主网络智能体的操作边界。

英文摘要

While LLM-based attackers exhibit growing proficiency in vulnerability exploitation, most existing cybersecurity benchmarks suffer from single-stage truncation, prematurely terminating evaluation upon initial access. In practice, initial footholds are exceptionally fragile across operational disruptions such as service restarts and host reboots. Whether LLM-based attackers can establish and maintain durable footholds beyond initial compromise remains a central blind spot in cybersecurity evaluation. To bridge this gap, we introduce CyberPersistBench, the first benchmark dedicated to post-compromise installation and persistence. Decoupled from upfront exploitation, CyberPersistBench frames persistence as an adversarial survival task in which agents use native host mechanisms to maintain footholds across staged system disruptions. Deterministic checks support a six-level scoring method (L1--L6) spanning installation and persistence. The benchmark comprises 203 core tasks across seven categories, augmented by multi-host and active defense extensions. Empirical evaluations across five frontier agents show that autonomous persistence remains limited (27.6%--44.8%) and drops further on defense-enabled tasks (5.5%--13.3%); nonetheless, these results reveal an emerging cyberattack risk, underscoring the necessity of benchmarking post-compromise persistence. CyberPersistBench thus establishes a foundational benchmark for post-compromise installation and persistence, delineating the operational boundaries of autonomous cyber agents.

Comments29 pages, 15 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑