发表机构
Fraunhofer FKIE; RWTH Aachen University(弗劳恩霍夫FKIE研究所; 亚琛工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文首次系统性评估基于风险的告警(RBA),将其重构为连续型告警优先级排序问题,通过8种数据集实验发现特定假设组合的RBA性能优于按告警严重程度排序的方法,可缓解网络安全告警疲劳。
AI 中文摘要
安全运营中心(SOC)面临大量虚假告警,在典型资源约束下难以检测网络攻击。基于风险的告警(RBA)被提出用于减少虚假告警,据报道在各类企业部署中已成功实现这一目标,但此前RBA尚未得到全面评估,其实施大多基于轶事证据,近乎猜测。本文对RBA开展首次系统性评估,为此将其重新表述为连续型告警优先级排序问题,而非二元决策问题(即是否超过告警阈值),从而可在所有可能阈值下评估性能,适配不同规模、不同告警量的SOC。研究提炼出5种基础风险假设,将其形式化为可独立参数化的模块,并在新型实验套件CATS中实现。针对8种不同告警数据集对假设进行全面评估,其中6种由本文创建或扩展以支撑此类评估。结果显示,特定假设组合实现了显著的告警优先级排序性能(8个数据集上的AUROC均值为0.92,标准差为0.09),优于按告警严重程度直接排序的方法(AUROC均值为0.72,标准差为0.21)。结论表明,RBA可大幅减少分析师需审查的虚假告警数量,具备缓解网络安全告警疲劳的潜力;此外,它还可为更复杂、资源密集型的告警分类方法(如基于大语言模型的方法)提供强基准。
英文摘要
Security operations centers (SOCs) face large numbers of false alerts, making detection of cyberattacks difficult under typical resource constraints. Risk-based alerting (RBA) has been proposed as a means to reduce false alerts and has reportedly succeeded in doing so in various enterprise deployments. However, RBA has not been comprehensively evaluated until now, leaving implementation mostly guesswork based on anecdotal evidence. In this paper, we present the first systematic evaluation of RBA. To this end, we reformulate it as a continuous alert prioritization problem rather than a binary decision problem (i.e., whether an alerting threshold is exceeded), allowing us to evaluate performance across all possible thresholds and thus model SOCs of varying sizes and alert volumes. We distill five fundamental risk hypotheses, formalize them as independently parametrizable modules, and implement them in our novel experimentation suite CATS. We thoroughly assess the hypotheses across eight diverse alert datasets, six of which we created or extended to make such an evaluation possible. Our results show that certain combinations of hypotheses achieve a remarkable alert prioritization performance (AUROC $μ=0.92$, $σ=0.09$ across the eight datasets), outperforming a straightforward prioritization by alert severity level (AUROC $μ=0.72$, $σ=0.21$). We conclude that RBA can substantially reduce the number of false alerts that analysts have to review and thus has the potential to mitigate cybersecurity alert fatigue. In addition, it serves as a strong baseline for more complex, resource-intensive alert triage approaches (e.g., based on large language models).
CommentsSubmitted to USENIX Security '27