发表机构
Indiana University; Rutgers University; Johns Hopkins University(印第安纳大学; 罗格斯大学; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究跨环境比较RL+RL与LLM+LLM两种分层红队架构,发现性能随环境反转,表明单一环境结论不可泛化,混合设计需针对特定失败模式。
AI 中文摘要
自主红队智能体通过规划策略和执行多阶段攻击,日益对AI赋能的网络防御进行压力测试。强化学习(RL)和大语言模型(LLMs)为这类智能体所需的规划与执行提供了互补机制,先前的工作已将它们以混合分层方式结合。然而,给定的架构通常是在单一环境中开发和评估的,这使得观察到的优势究竟反映了一种普遍更强的决策机制,还是仅仅与特定设置相匹配,仍悬而未决。我们通过受控的跨环境比较来填补这一空白,比较两种同构分层红队架构:RL规划器配RL执行器(RL+RL)和LLM规划器配LLM执行器(LLM+LLM)。我们在CybORG CAGE-4和两种网络规模下的Cyberwheel中,针对专家自主防御者评估这两种架构,共覆盖18种配置,并采用统一的破坏性指标。我们发现了一种显著的环境依赖性反转。RL+RL在紧凑、奖励密集的CAGE-4中获胜(78.5%的破坏成功率,而最强LLM配置为18.0%),并在100主机Cyberwheel网络中获胜(81.0%对50.5%);而一个预训练的网络安全LLM智能体在更大、具有升级门槛的1010主机Cyberwheel网络中获胜(55.0%对RL的0.0%)。杀伤链分析通过架构特定的瓶颈解释了这种反转,这些瓶颈聚合了成功率(见http URL)。在1010主机Cyberwheel网络中,RL发现并攻陷主机,但在权限提升阶段停滞;而在CAGE-4中,LLM智能体获得特权访问,但很少将其转化为实际影响。这些结果表明,在单一环境中得出的结论可能无法泛化,混合规划器-执行器设计应基于特定的失败模式来驱动,而非假设某一种架构普遍更优。
英文摘要
Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning (RL) and large language models (LLMs) offer complementary mechanisms for the planning and execution such agents require, and prior work has combined them in hybrid hierarchies. Yet a given architecture is typically developed and evaluated within a single environment, leaving open whether an observed advantage reflects a generally stronger decision mechanism or merely alignment with a particular setting. We address this gap with a controlled cross-environment comparison of two homogeneous hierarchical red team architectures: an RL planner with an RL executor (RL+RL) and an LLM planner with an LLM executor (LLM+LLM). We evaluate both against expert autonomous defenders in CybORG CAGE-4 and in Cyberwheel at two network scales, across 18 configurations under one unified disruption metric. We find a pronounced environment-dependent inversion. RL+RL wins the compact, densely rewarded CAGE-4 (78.5% disruption success versus 18.0% for the strongest LLM configuration) and the 100-host Cyberwheel network (81.0% versus 50.5%), while a pretrained cybersecurity LLM agent wins the larger, escalation-gated 1010-host Cyberwheel network (55.0% versus 0.0% for RL). A kill-chain analysis explains the inversion through architecture-specific bottlenecks that aggregate success rates conceal.In the 1010-host Cyberwheel network, RL discovers and compromises hosts but stalls at privilege escalation, whereas in CAGE-4, LLM agents obtain privileged access but rarely convert it into operational impact. These results indicate that conclusions drawn in a single environment may not generalize, and that hybrid planner-executor designs should be motivated by specific failure modes rather than the assumption that one architecture is universally preferable.