迈向基于大语言模型的层次化网络防御:从规划到执行
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
- Indiana University(印第安纳大学)
- Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究探讨冻结零样本大语言模型在层次化网络防御中从规划到执行的控制能力,发现能力足够强的70B模型无需重新训练即可在多种网络规模下保持强防御性能,且战术执行对发挥LLM优势至关重要。
AI中文摘要:
使用强化学习(RL)训练的自主任网络防御者通常与其训练时所处的网络绑定,这限制了其在网络规模变化时的泛化能力。层次化强化学习通过将战略目标选择与战术执行分离来降低决策复杂性,但并未消除这种重新训练依赖性。我们研究了冻结的、零样本的大语言模型(LLMs)能否在层次化网络防御中提供无需重新训练的控制,以及当LLM控制从规划扩展到执行时性能如何变化。我们构建了一个控制器无关的规划器-执行器层次结构,其中规划器在固定时间范围内选择要防御的子网,执行器在该子网内选择防御动作。利用高保真的Cyberwheel环境(其内置的自动化红队代理映射到MITRE ATT&CK框架),我们比较了RL+RL、LLM+RL和LLM+LLM三种配置,使用了从3B到70B参数的六种模型,包括两种网络安全专用模型,并在小型、中型和大型网络上进行了评估。仅将规划器替换为LLM在网络规模增大时带来的收益有限。相比之下,将LLM控制扩展到执行阶段,对于能力足够强的模型产生了显著改进。例如,一个冻结的通用70B模型在所有三种网络规模下,使用相同的模型权重,将成功的横向移动控制在约1%的步骤内,并将攻击者影响降至接近零,而RL基线则需针对每种规模重新训练。我们的结果表明,能力足够强的冻结LLM可以在无需任务特定重新训练的情况下,在评估的网络规模上保持强大的防御性能,同时也表明强大的战术执行对于实现基于LLM的控制优势至关重要。
英文摘要:
An autonomous cyber defender trained with reinforcement learning (RL) is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it does not eliminate this retraining dependence. We investigate whether frozen, zero-shot large language models (LLMs) can provide retraining-free control in hierarchical cyber defense and how performance changes as LLM control is extended from planning to execution. We formulate a controller-agnostic planner-executor hierarchy in which the planner selects a subnet to defend over a fixed horizon and the executor selects defensive actions within that subnet. Using the high fidelity Cyberwheel environment, with its built-in automated red team agent mapped to the MITRE ATT&CK framework, we compare RL+RL, LLM+RL, and LLM+LLM configurations using six models ranging from 3B to 70B parameters, including two cybersecurity-specialized models, across small, medium, and large networks. Replacing only the planner with an LLM yields limited gains as network size increases. In contrast, extending LLM control to execution produces notable improvements for sufficiently capable models. For instance, a frozen general purpose 70B model holds successful lateral movement to approximately 1% of steps and attacker impact near zero across all three network scales using the same model weights, while the RL baseline is retrained for each scale. Our results show that sufficiently capable frozen LLMs can maintain strong defensive performance across the evaluated network scales without task-specific retraining, while also indicating that strong tactical execution is important to realizing the benefits of LLM-based control.