发表机构
The University of Texas at El Paso(德克萨斯大学埃尔帕索分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一个将状态对齐与动作转换分离的框架,实现网络攻击智能体在多个模拟器间及模拟器到真实环境的零样本策略迁移,无需重新训练,并在实验中验证了其可行性与行为相似性。
AI 中文摘要
网络攻击智能体通常在单一模拟器中进行训练和评估,这使得学习到的策略能否迁移到开发环境之外变得不明确。这一局限性阻碍了部署和公平比较,因为网络模拟器在状态表示、观测模型和动作空间上存在显著差异。本文研究了跨网络环境的策略迁移,并论证了模拟器到模拟器以及模拟器到真实环境的迁移可以视为同一底层对齐问题的实例。我们提出了一个将状态对齐与动作转换分离的框架,使得在一个环境中训练的策略无需重新训练即可在另一个环境中运行。我们在四个网络平台(CyberBattleSim、NetSecGame、CyberWheel和NASim)上评估了迁移效果,包括在NASim中的仿真部署。实验表明,零样本迁移是可行的,在紧密对齐的环境中完全保留了源策略的性能,并在将源性能为60.5%的策略迁移时实现了45.2%的胜率。在仿真的虚拟机环境中,迁移策略与原生策略的Jensen-Shannon散度为0.085,表明行为相似性很强。代码和基准可在以下网址获取:this https URL。
英文摘要
Cyber attack agents are typically trained and evaluated within a single simulator, making it unclear whether learned policies transfer beyond the environments in which they were developed. This limitation hinders both deployment and fair comparison, as cyber simulators differ substantially in their state representations, observation models, and action spaces. In this paper, we study policy transfer across cyber environments and argue that simulator-to-simulator and simulator-to-real transfer can be viewed as instances of the same underlying alignment problem. We propose a framework that separates state alignment from action translation, enabling a policy trained in one environment to operate in another without retraining. We evaluate transfer across four cyber platforms, CyberBattleSim, NetSecGame, CyberWheel, and NASim, including emulated deployments in NASim. Our experiments show that zero-shot transfer is feasible, fully preserving source-policy performance in closely aligned environments and achieving 45.2% win rates when transferring policies whose source performance is 60.5%. In emulated virtual machine environments, transferred policies exhibit a Jensen-Shannon divergence of 0.085 from native policies, indicating strong behavioral similarity. Code and benchmarks are available at: https://anonymous.4open.science/r/RL-Transfer-between-env-4F47/.
Comments13 pages, 3 figures, 1st Workshop on Real-world AI Security and Engineering for Cybersecurity Systems (RAISE) 2026