发表机构
University of California, Irvine; Northeastern University; Johns Hopkins University(加州大学欧文分校; 东北大学; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对DRL网络防御对自适应威胁鲁棒性不足的问题,提出Trident智能体式LLM红队框架,通过含三类组件的设计,发现现有防御存在根本性脆弱性,大幅降低蓝方防御性能并揭示新兴行为。
AI 中文摘要
基于深度强化学习(DRL)的自主网络防御系统已受到大量研究关注,但几乎仅针对静态启发式红方智能体进行评估,其对自适应威胁的鲁棒性研究严重不足。同时,近期可验证奖励强化学习(RLVR)的进展提升了大语言模型(LLM)的推理能力,但由于缺乏合适的基准环境和交互数据集,其在网络安全领域的应用仍难以实现。为填补这一空白,我们提出Trident(三叉戟),一种智能体式LLM红队测试框架,包含三个组件:一是涵盖CybORG CAGE 4和CyberWheel的隔离沙箱服务器动态基准;二是包含超过13000条用于RLVR的高保真红蓝交互轨迹的数据集;三是“代码即策略”RLVR智能体架构(Trident Agentic)。后者将红方智能体训练通过三方“日志摘要器-规划器-编码器”设计重新表述为上下文多臂老虎机问题,其中可训练的规划器从压缩的执行日志生成完整攻击策略,冻结的编码器将其转换为可执行的Python策略,用于对抗实时DRL防御者。实证评估显示现有防御存在根本性脆弱性:使用单个可训练的7B规划器,Trident与静态红方智能体基线相比平均降低蓝方智能体防御性能522%,同时自主发现静态启发式完全无法揭示的新兴行为,如诱饵规避和自适应状态优先级排序。
英文摘要
Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost exclusively against static, heuristic red agents, leaving their robustness against adaptive threats critically understudied. Meanwhile, recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have improved LLM reasoning, but their integration into cybersecurity remains elusive due to the absence of suitable benchmark environments and interaction datasets. To bridge this gap, we introduce Trident, an agentic LLM red teaming framework comprising three components: a dynamic benchmark with isolated sandbox servers spanning CybORG CAGE 4 and CyberWheel, a dataset comprises over 13,000 high-fidelity red-blue interaction trajectories for RLVR, and a ``Code-as-Policy'' RLVR agentic architecture Trident Agentic). The latter reformulates red agent training as a contextual bandit via a tripartite Log Summarizer--Planner--Coder design, where a trainable Planner generates complete attack strategies from compressed execution logs, which a frozen Coder translates into executable Python policies deployed against live DRL defenders. Empirical evaluations reveal a fundamental brittleness in existing defenses: with a single trainable 7B planner, Trident reduces blue agent defensive performance by an average of 522% compared to static red agent baselines while autonomously discovering emergent behaviors such as decoy avoidance and adaptive state prioritization that static heuristics entirely fail to uncover.
Commentscode: https://github.com/BiasLabProjects/Trident