arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

请教专家:基于LLM引导的强化学习用于自主网络防御

Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense

Fernando Martinez, Abhishek Satyam, Tao Li, Junaid Farooq, Ying Wang, Juntao Chen

arXiv 2610.09337首次发表:更新:

发表机构

Fordham University; City University of Hong Kong; University of Michigan-Dearborn; Stevens Institute of Technology(福特汉姆大学; 香港城市大学; 密歇根大学迪尔伯恩分校; 史蒂文斯理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Ask the Expert框架,利用LLM在训练时提供防御建议并转化为分层奖励塑形,提升PPO在自主网络防御中的样本效率,部署时无需LLM依赖。

AI 中文摘要

基于策略的强化学习(RL)方法在自主网络防御中已产生有前景的结果;然而,在防御者必须面对延迟、部分观测以及大动作空间的情况下做出响应时,这些方法样本效率低下。虽然大型语言模型(LLMs)可以对安全状态空间进行语义推理,但高延迟和信任假设阻碍了有吸引力的在线部署模型。我们提出了“请教专家”(Ask the Expert),一种训练时引导框架,该框架首先总结困难的网络防御状态,然后通过受限动作接口间歇性地查询LLM以获得主机级防御建议,最后将这些建议转化为分层奖励塑形(tiered reward shaping)以供PPO使用。由于LLM在训练后被丢弃,部署时是纯RL策略。在TTCP CAGE CC1和CC2以及两种攻击者类型上,这种非对称设计相比PPO提高了样本效率,并优于所评估的基于势能的奖励塑形(PBRS)基线,同时保持了最强的终端均值,且在部署时无需LLM依赖。

英文摘要

Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semantically about security state space, high latency and trust assumptions prevent attractive in-line deployment models. We introduce Ask the Expert, a training-time guidance framework which first summarizes hard cyber-defense states, then intermittently queries an LLM for host-level defensive recommendations via a constrained action interface, and finally transforms those recommendations into tiered reward shaping for use with PPO. Because the LLM is discarded after training, deployment is a pure RL policy. Across TTCP CAGE CC1 and CC2 and both attacker types, this asymmetric design improves sample efficiency over PPO and outperforms the evaluated potential-based reward shaping (PBRS) baselines, while retaining the strongest terminal mean and requiring no LLM dependency at deployment time.

CommentsAccepted for publication at IEEE GLOBECOM 2026. Proceedings forthcoming. 6 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑