AI 中文总结
本文通过对20位不同类型SOC从业者的半结构化访谈,分析了LLM在SOC中的应用现状、存在的局限性,提出了LLM集成成熟度标准及相关研究议程。
AI 中文摘要
大型语言模型(LLM)正越来越多地被应用于安全运营中心(SOC),以支持警报情境化、事件总结和调查文档起草等文本密集型分析工作。尽管人们对此兴趣浓厚,但从业者指出了关键的运营问题,最显著的是幻觉(看似合理但实则错误的输出)、不透明的推理,以及在安全工作流中安全使用模型生成内容所需的验证工作。在本文中,我们呈现了对20位SOC从业者进行半结构化访谈的结果,这些从业者涵盖一线分析师、SOC经理和工具开发者。参与者表示,在可快速验证的低风险任务(如日志总结或起草初步调查线索)上能感知到时间节省,但他们始终将LLM输出视为初步草稿和建议,而非决策级结论。参与者还表示,由于输出不可靠且模型推理不明确,他们对LLM用于高风险安全决策的信任有限,并报告称主要依赖临时验证规范和持续的人工监督,而非标准化缓解程序。基于这些访谈得出的描述,我们引入了一个成熟度标准来表征LLM集成的准备情况,并概述了一项研究议程,重点关注可审计性和透明解释机制,以支持在SOC工作流中更安全地采用LLM。
英文摘要
Large Language Models (LLMs) are increasingly being explored within Security Operation Centers (SOCs) to support text-heavy analytical work such as alert contextualization, incident summarization, and drafting investigative artifacts. Despite this interest, practitioners describe critical operational concerns, most notably hallucinations (plausible but incorrect outputs), opaque reasoning, and the verification effort required to safely use model-generated content in security workflows. In this paper, we present findings from semi-structured interviews with 20 SOC practitioners spanning frontline analysts, SOC managers, and tool developers. Participants report perceived time savings for low-stakes tasks that are quickly verifiable (e.g., summarizing logs or drafting initial investigative leads), but they consistently frame LLM outputs as preliminary drafts and suggestions rather than decision-grade conclusions. Participants also describe limited trust in LLMs for high-stakes security decisions due to unreliable outputs and unclear model reasoning, and they report relying primarily on ad-hoc verification norms and continuous human oversight rather than standardized mitigation procedures. Based on these interview-grounded accounts, we introduce a maturity rubric to characterize readiness for LLM integration and outline a research agenda emphasizing auditability and transparent explanation mechanisms to support safer adoption in SOC workflows.