arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从噪声到信号:利用具备端点特定日志的大语言模型改进安全日志异常检测

From Noise to Signal: Improving Security Log Anomaly Detection Using LLMs with Endpoint-Specific Logs

Christopher Henshaw, Gour Karmakar

arXiv 2608.19938首次发表:更新:

发表机构

Institute of Innovation, Science and Sustainability; Centre for Smart Analytics; Federation University Australia(创新、科学与可持续发展研究所; 智能分析中心; 澳大利亚联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发基于指令的LLM分类框架,结合端点特定日志,经实验验证Meta Llama 3.1 8B Instruct在安全日志异常检测中性能优于Wazuh、OpenSearch等方法。

AI 中文摘要

现有的异常行为日志检测方法(如Wazuh)主要依赖预定义的检测规则,而基于统计的异常检测方法(如OpenSearch)则识别与先前观测到的行为模式的偏差。近期研究已探索将大语言模型(LLMs)用于日志异常检测,因其具备解释语义和上下文信息的能力。然而,基于LLM的方法可能受提示工程、噪声日志数据以及对可能缺乏端点特定认证行为的通用数据集的依赖影响。为解决这些局限,本研究开发了一种标准化的基于指令的LLM分类框架,用于检测异常认证行为,包括边界案例。研究构建了受控网络安全测试平台以生成端点特定的认证数据,产出了包含正常、边界及异常行为场景的精选数据集。本研究采用通用真实值严重程度框架,将三个经指令微调的LLMs——Meta Llama 3.1 8B Instruct、Qwen 2.5 7B Instruct和GPT-OSS 20B——与Wazuh基于规则的检测及OpenSearch异常检测进行评估。Meta Llama 3.1 8B Instruct实现了最强的整体端到端检测性能,准确率达89.3%、召回率为88.2%、F1分数为91.8%、漏报率为11.8%;相比之下,Wazuh准确率为52.0%、漏报率为68.6%,OpenSearch准确率为49.3%、漏报率为74.5%。Meta Llama还检测到80%的边界异常场景,而Wazuh和OpenSearch分别仅检测到20%和15%。Qwen的整体检测性能低于Meta Llama,但平均推理延迟最低,且结构化响应有效性达100%;GPT-OSS在生成有效响应时展现出强劲的分类性能。

英文摘要

Existing approaches to anomalous behaviour log detection, such as Wazuh rely primarily on predefined detection rules, while statistical anomaly detection approaches such as OpenSearch identify deviations from previously observed behavioural patterns. Recent research has investigated LLMs for log anomaly detection because of their ability to interpret semantic and contextual information. However, LLM-based approaches can be affected by prompt construction, noisy log data, and reliance on generic datasets that may lack endpoint-specific authentication behaviours. To address these limitations, this study develops a standardised instruction-based LLM classification framework for detecting anomalous authentication behaviours, including borderline cases. A controlled cybersecurity testbed was developed to generate endpoint-specific authentication data, producing a curated dataset comprising normal, borderline, and anomalous behavioural scenarios. Three instruction-tuned LLMs, Meta Llama 3.1 8B Instruct, Qwen 2.5 7B Instruct, and GPT-OSS 20B, were evaluated against Wazuh rule-based detection and OpenSearch Anomaly Detection using a common ground-truth severity framework. Meta Llama 3.1 8B Instruct achieved the strongest overall end-to-end detection performance, with an accuracy of 89.3%, recall of 88.2%, F1-score of 91.8%, and false negative rate of 11.8%. In comparison, Wazuh achieved an accuracy of 52.0% and false negative rate of 68.6%, while OpenSearch achieved an accuracy of 49.3% and false negative rate of 74.5%. Meta Llama also detected 80% of the borderline anomalous scenarios, compared with 20% for Wazuh and 15% for OpenSearch. Qwen achieved lower overall detection performance than Meta Llama but recorded the lowest average inference latency and 100% structured-response validity. GPT-OSS demonstrated strong classification performance when valid responses were produced.

CommentsThe paper contains 35 pages and 3 figures. The paper has not been submitted or published in any conference or journal. The authors have an aim to publish it in a journal

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑