arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15012cs.CRcs.AIcs.MA

SysEvolve:一种原生AI、安全且自主的对抗攻防协同进化系统

SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system

Yuhan Meng, Shaofei Li, Jionghao Huang, Jiandong Jin, Puyi Wang, Hanlin Jiang, Anis Yusof, Peng Jiang, Zhenkai Liang, Yao Guo, Ding Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对网络安全攻防不对称问题,提出SysEvolve协同进化系统,含三个组件,实验验证其攻防性能,还揭示了LLM智能体能力的三项关键发现。

中文摘要 AI 辅助

大型语言模型(LLM)的快速发展造成了网络安全领域日益加剧的不对称性:攻击正朝着自主执行的方向加速,而防御仍主要依赖人力密集型工作。尽管此前在网络靶场、AI驱动攻击和AI驱动防御方面已开展了大量研究,但这种不对称性依然存在。我们将其根源追溯至攻防双方的进化本身在三个层面陷入停滞。为克服这一问题,我们提出协同进化作为整合思路,即攻防AI智能体通过对抗性对抗自主且安全地推动彼此的进化。基于这一思路,我们提出了SysEvolve,它包含三个协同设计的组件:SysField、SysSpear和SysArmor。SysField构建逼真的多主机靶场;SysSpear生成高效、安全的攻击方案;SysArmor执行实时、可解释的防御。三者共同构成一个自驱动的对抗循环,恢复所有三个层面的进化。在评估中,SysField实现了零损失收集,仅产生2.1%的开销,并将257个CVE编排为1148个靶场;SysSpear相比基线LLM将攻击成功率提升了25%以上;SysArmor的精度比现有系统高10至1000倍,且在华为和深信服的生产环境中检测到了真实的APT攻击。我们的评估还揭示了关于LLM智能体能力的三项发现:第一,多步骤组合和更大的拓扑结构暴露了单步骤评估所隐藏的智能体能力差距;第二,瓶颈在于初始访问之后的入侵后状态利用阶段;第三,LLM智能体易受环境干扰:当在靶场中部署诱饵端点时,尽管初始访问成功率保持不变,但智能体的超时时间增加了两倍,且下游完成情况消失。

英文摘要

The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward autonomous execution while defense remains predominantly human-intensive. Despite substantial prior work across cyber ranges, AI-driven attack, and AI-driven defense, this asymmetry persists. We trace it to a deeper root cause, that evolution itself has stalled on both sides at three layers. To overcome this, we propose co-evolution as the integrating insight, where attack and defense AI agents autonomously and safely drive each other's evolution through adversarial confrontation. Based on this insight, we present \sysevolve, comprising three co-designed components, \sysfield, \sysspear, and \sysarmor. \sysfield constructs realistic multi-host ranges. \sysspear generates efficient, safe attack schemes. \sysarmor performs real-time, interpretable defense. Together they form a self-driven adversarial loop restoring evolution at all three layers. In evaluation, \sysfield achieves zero-loss collection at 2.1\% overhead and orchestrates 257 CVEs into 1,148 ranges, \sysspear improves attack success by over 25\% over baseline LLMs, and \sysarmor achieves 10--1000$\times$ greater precision than prior systems and detects real APT attacks in production at Huawei and Sangfor. Our evaluation also reveals three findings about LLM agent capabilities. First, multi-step composition and larger topologies expose agent capability gaps hidden by single-step evaluations. Second, the bottleneck lies after initial access in post-compromise state utilization. Third, LLM agents are susceptible to environmental interference. When decoy endpoints are deployed in the range, agent timeouts triple and downstream completion disappears despite the success rates of initial accesses are unchanged.

发表机构

  • School of Computer Science, Peking University(北京大学计算机学院)
  • Southeast University(东南大学)
  • School of Computing, National University of Singapore(新加坡国立大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑