AI 中文总结
该研究采用SEBoK指导的流程,将安全目标、利益相关者需求等作为提示输入5款LLMs,生成并整合27条安全需求,经专家评估后证实LLMs可辅助安全需求提取早期阶段,但需人工审核优化。
AI 中文摘要
网络靶场是包含众多交互组件和具有不同安全关注点的利益相关者的复杂环境,面向服务的网络靶场(SOR)也不例外,尤其是在针对关键基础设施的培训场景中。安全关注点被转化为安全需求,提取这些需求通常既困难又耗时。本研究探讨大语言模型如何协助提取面向服务的网络靶场的安全需求,并为设计者和开发者生成有用的基线。该方法遵循由系统工程与软件工程知识体系(SEBoK)指导的流程:首先识别安全使命目标和利益相关者需求,随后将这些内容与架构指南一起作为提示上下文提供给5个LLMs,分别是GPT-5.2、Gemini 3.1 Pro、Grok 4.1、Sonar和Kimi K2.5。这些模型共生成84条安全需求,经整合后形成包含27条需求的综合集合,并映射到面向服务的网络靶场的架构层。最终集合由5名网络安全专家按照必要性、清晰性、完整性、可行性、可测试性标准及额外的弃权(不执行)选项进行评估。结果显示接受率较高:必要性98.5%、清晰性87.4%、完整性85.2%、可行性78.5%、弃权0.7%;可测试性较低,为44.4%,表明关于这些需求如何测试的信息略有不足。研究结果表明,LLMs可支持安全需求提取的早期阶段,但仍需人工审核,尤其是为了改进或调整需求的某些方面。
英文摘要
Cyber ranges are complex environments comprising many interacting components and stakeholders with different security concerns. The Service-Oriented Cyber Range (SOR) is no exception, particularly when it comes to training scenarios targeting critical infrastructure. Security concerns are translated into security requirements, the elicitation of which is usually difficult and time-consuming. This work examines how large language models can assist in eliciting security requirements for a service-oriented range and help produce a useful baseline for designers and developers. The approach follows a SEBoK-guided process in which security mission objectives and stakeholder needs were first identified and then provided as a prompt context along with architectural guidelines to five LLMs: GPT-5.2, Gemini 3.1 Pro, Grok 4.1, Sonar, and Kimi K2.5. The models generated 84 security requirements in total, which were consolidated into a comprehensive set of 27 requirements and then mapped to the architectural layers of the service-oriented range. The final set was evaluated by five cybersecurity experts against the criteria of necessity, clarity, completeness, feasibility, and testability, with an additional rejection option. The results showed a high acceptance rate, specifically for necessity with 98.5%, clarity with 87.4%, completeness with 85.2%, feasibility with 78.5%, and rejection with 0.7%. Testability was lower at 44.4%, indicating a slight lack of information on how these requirements could be tested. These findings show that LLMs can support early stages of the elicitation of security requirements, although human review is still needed, especially to improve or adjust certain aspects of the requirements.