arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28826cs.LG

将大型语言模型中的知识蒸馏为用于自主网络操作的轻量级强化学习智能体

Distilling Knowledge from Large Language Models into Lightweight Reinforcement Learning Agents for Autonomous Cyber Operations

Konur Tholl, François Rivest, Mariam El Mezouar, Adrian Taylor, Ranwa Al Mallah

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对自主网络操作中强化学习智能体训练不稳定的问题,通过提示工程证明80亿参数网络安全LLM性能优于基线RL智能体,再经在线策略蒸馏将其防御知识迁移至仅64910参数的轻量级RL智能体,为前沿网络安全模型的实用化提供了可行方案。

中文摘要 AI 辅助

自主网络操作(Autonomous Cyber Operations, ACO)对于防御企业网络愈发重要,因为网络威胁的复杂性不断提升。ACO应用通常采用强化学习(Reinforcement Learning, RL)智能体,通过与环境交互学习防御行为。然而,RL智能体在训练时通常需要大量探索,在收敛到有效防御策略前,往往会出现行为不稳定、初始决策质量差的问题。本研究探索利用大型语言模型(Large Language Model, LLM)提升ACO环境中的自主防御决策。我们采用提示工程而非微调,证明一个在网络安全数据上预训练的80亿参数LLM,在修改后的CybORG CAGE Challenge 2环境中,性能优于基线RL智能体。随后,我们提出一种在线策略蒸馏框架,将LLM的防御策略迁移至仅含64910个参数的轻量级RL智能体,使模型规模缩减数个数量级,同时保持有效防御能力,为将前沿网络安全模型部署到轻量级、可部署智能体提供了途径。为评估可迁移性,我们构建了主机数量从4到12的CybORG场景,并在不同网络配置下评估该方法。我们还评估了教师引导的RL稳定策略,发现没有一种策略能始终超越优化后的教师策略,这表明奖励驱动的RL优化与教师引导的防御策略之间存在策略对齐的局限性。我们的结果证明,以网络安全为重点的LLM有潜力作为自主网络防御的专业知识来源,而策略蒸馏为将前沿网络安全模型部署到高效、可扩展的智能体提供了实用路径。

英文摘要

Autonomous Cyber Operations (ACO) are increasingly important for defending enterprise networks as cyber threats continue to evolve in sophistication. ACO applications commonly employ Reinforcement Learning (RL) agents to learn defensive behaviors through interaction with environments. However, RL agents typically require extensive exploration during training, often resulting in unstable behavior and poor initial decision-making before converging toward effective defense strategies. In this work, we investigate the use of a Large Language Model (LLM) to improve autonomous defensive decision-making within an ACO environment. Through prompt engineering rather than fine-tuning, we demonstrate that an 8-billion parameter LLM pretrained on cybersecurity data can outperform a baseline RL agent in a modified CybORG CAGE Challenge 2 environment. We then propose an online policy distillation framework that transfers the LLM's defensive policy into a lightweight RL agent containing only 64,910 parameters, reducing model size by several orders of magnitude while maintaining effective defensive capabilities. This provides a pathway toward operationalizing frontier cybersecurity models within lightweight, deployable agents. To evaluate transferability, we construct CybORG scenarios ranging from 4 to 12 hosts and assess the approach across varying network configurations. We also evaluate teacher-guided RL stabilization strategies and observe that none consistently surpass the optimized teacher policy, suggesting policy-alignment limitations between reward-driven RL optimization and teacher-guided defense strategies. Our results demonstrate the potential of cybersecurity-focused LLMs as sources of expertise for autonomous cyber defense, while policy distillation provides a practical path toward operationalizing frontier cybersecurity models within efficient, scalable agents.

发表机构

  • Royal Military College of Canada(加拿大皇家军事学院)
  • Defence Research and Development Canada(加拿大国防研究与发展部)
  • Polytechnique Montreal(蒙特利尔理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑