发表机构
Virginia Tech; Google; North Carolina State University(弗吉尼亚理工大学; 谷歌; 北卡罗来纳州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM智能体在SOC警报分类中漏报率高的问题,构建ALERT-BENCH基准测试五种方法,提出多智能体框架AIDA,使漏报率降至3.1%,F1分数达0.958,显著提升了SOC警报分类性能。
AI 中文摘要
安全运营中心(SOC)必须对大量警报进行分类,其中大部分是良性的,而未被发现的攻击可能会被搁置。使用工具的大语言模型(LLM)智能体可以在分类过程中检索证据,但推理策略如何决定收集什么以及何时调查足以关闭警报仍不清楚。我们研究了五种代表性方法,涵盖单次工具使用、迭代检索、抽样调查、自我审查和显式验证。为支持这项研究,我们构建了ALERT-BENCH,这是一个交互式基准,通过实时安全信息和事件管理系统(SIEM)重放企业遥测数据,并要求每个系统检索证据。在来自多阶段攻击场景的1247个警报中,每种方法都遗漏了至少40.4%的与攻击相关的警报。轨迹分析显示,当搜索返回无记录时,攻击警报更有可能被驳回;同上下文审查的净修正为负;驳回所获得的调查力度并不比升级更强。基于这些发现,我们进一步设计了AIDA(对抗性调查与辩证分析),这是一个多智能体框架,要求在独立质疑前提出明确的决策建议,且在驳回前有更强的证据要求。AIDA在仅追加的调查账本中保留调查历史,并将质疑保留在单独的推理上下文中。一个独立的裁决者根据证据对决策建议和质疑进行裁决,当证据缺失时,要么解决警报,要么请求另一轮调查。在相同的警报上,AIDA的F1分数为0.958,而所研究方法的F1分数为0.371-0.744;它将漏报率从40.4%降至3.1%,同时将18.4%的警报升级给分析师。这些结果表明,构建证据检索和决策审查的结构可以显著提升智能体式SOC分类的性能。
英文摘要
Security operations centers (SOCs) must triage large volumes of alerts, most of which are benign, while missed attacks can remain uninvestigated. Tool-using large language model (LLM) agents can retrieve evidence during triage, but it remains unclear how reasoning strategies determine what to gather and when an investigation is sufficient to close an alert. We study five representative approaches spanning single-pass tool use, iterative retrieval, sampled investigations, self-review, and explicit verification. To support this study, we build ALERT-BENCH, an interactive benchmark that replays enterprise telemetry through a live SIEM and requires each system to retrieve evidence. Across 1,247 alerts from a multi-stage attack scenario, every approach missed at least 40.4% of attack-related alerts. Trace analysis shows that attack alerts are more likely to be dismissed when searches return no records, same-context review has negative net correction, and dismissal receives no consistently stronger investigation than escalation. Based on these findings, we further design AIDA (Adversarial Investigation and Dialectical Analysis), a multi-agent framework that requires an explicit proposed decision before independent challenge and stronger evidentiary requirements before dismissal. AIDA preserves investigation history in an append-only Investigation Ledger and keeps the challenge in a separate reasoning context. A separate Judge adjudicates the proposed decision and challenge against evidence, resolving the alert or requesting another round when evidence is missing. On the same alerts, AIDA achieves an F1 score of 0.958, compared with 0.371-0.744 for the studied approaches, and reduces the false-negative rate from 40.4% to 3.1% while escalating 18.4% of alerts to analysts. These results show that structuring evidence retrieval and decision review can substantially improve agentic SOC triage.
CommentsPreprint