arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SENTINEL-RL:安全运营中心中大型语言模型智能体的拓扑推理卸载方案

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild

arXiv 2609.04159首次发表:更新:

发表机构

Luddy School of Informatics, Computing and Engineering; Indiana University(勒迪信息学、计算与工程学院; 印第安纳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Sentinel-RL将SOC中LLM智能体的拓扑推理解耦,通过PPO策略等技术实现高效的认证图处理与安全操作,在数据集和HPC集群上取得多项性能成果并贡献了相关工程与部署模式。

AI 中文摘要

大型语言模型(LLM)智能体正被越来越多地提议作为自主安全运营中心(SOC)分析师,但企业级应用存在两大局限:有限的上下文窗口无法容纳数千台主机的认证图,且自由生成模式无法保证推荐的遏制操作与其作用的拓扑结构一致。我们提出Sentinel-RL,一种将拓扑推理与语义推理解耦的智能SOC架构:异构图注意力编码器将实时认证子图总结为固定维度的状态,近端策略优化(PPO)策略将该状态映射到一组受约束的调查操作,LLM智能体循环仅负责消费策略的推荐并生成由评判器控制的分析师可读叙述。我们在LANL多源网络安全事件综合数据集和印第安纳大学Quartz高性能计算(HPC)集群上实例化该系统,报告四项结果:(i)两阶段CREATE摄取模式在单个32核节点上耗时14.2分钟将含2400万条边的认证子图加载到Neo4j,比基于标准MERGE的管道快约24倍;(ii)滑动窗口警报引擎在50次试验中均能在≤2.5秒内触发25事件/10秒的阈值;(iii)PPO训练经200次迭代收敛到平均回合回报8.74±0.31,在保留的标记红队事件上的精度为0.91、召回率为0.87;(iv)集成遏制循环完成完整的检测-调查-推荐-人工批准周期的中位数耗时6.3秒。我们贡献了可复用的工程模式(热节点死锁规避方案)、可移植的HPC部署模式(锚节点共置),以及涵盖误报经济性、可逆性保证、审计合规性和人工批准边界的企业就绪性分析。

英文摘要

Large language model (LLM) agents are increasingly proposed as autonomous SOC analysts, but two limitations make them unreliable at enterprise scale: a finite context window cannot hold a multi-thousand-host authentication graph, and free-form generation offers no guarantee that a recommended containment action is consistent with the topology it operates on. We present Sentinel-RL, an agentic-SOC architecture that decouples topological reasoning from semantic reasoning: a heterogeneous graph attention encoder summarizes the live authentication subgraph into a fixed-dimensional state, a Proximal Policy Optimization (PPO) policy maps this state to a constrained set of investigative actions, and an LLM agent loop is restricted to consuming the policy's recommendations and producing analyst-readable narratives gated by a critic. We instantiate the system on the LANL Comprehensive, Multi-Source Cyber-Security Events dataset and the Indiana University Quartz HPC cluster, reporting four results: (i) a two-phase CREATE ingestion pattern loads a 24M-edge authentication subgraph into Neo4j in 14.2 minutes on a single 32-core node, roughly 24x faster than the canonical MERGE-based pipeline; (ii) a sliding-window alert engine reliably trips a 25-event/10-second threshold in <=2.5 s across 50 trials; (iii) PPO training over 200 iterations converges to a mean episodic return of 8.74+/-0.31, with held-out precision of 0.91 and recall of 0.87 on labeled red-team events; and (iv) the integrated containment loop completes a full detect-investigate-recommend-human-approve cycle in a median of 6.3 s. We contribute a reusable engineering pattern (the hot-node deadlock workaround), a portable HPC deployment pattern (anchor-node co-location), and an enterprise-readiness analysis covering false-positive economics, reversibility guarantees, audit compliance, and the human-approval boundary.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑