AI 中文总结
研究针对野外攻击调查中依赖爆炸和因果链碎片化问题。提出SherAgent系统,基于大语言模型,采用迭代“查询-过滤”回溯范式处理非结构化数据,克服相关问题。相比传统方法,提高调查成功率,且运行高效,减轻专家分析负担。
AI 中文摘要
基于溯源的攻击调查通过标准化数据和查询逻辑实现可行的自动化;然而,在实际中它受到野外依赖爆炸和因果链碎片化的严重阻碍。为设计一个强大的自动化调查工具,我们与一家服务数十亿用户的大型互联网公司的安全运营中心合作。通过参与实际事件响应,评估并改进其现有的基于大语言模型的调查工作流程,找出调查失败的根本原因和现有工具中的主要挑战。受这些发现启发,我们提出SherAgent,一个大语言模型赋能的自动化调查系统。它在溯源图上采用迭代的“查询-过滤”回溯范式,利用大语言模型的语义推理能力处理非结构化数据,如调查上下文和威胁情报。为克服由缺失事件导致的因果链碎片化,系统动态校准查询条件以扩大搜索范围,同时进行精确结果过滤和战略节点选择以减轻依赖爆炸。野外的广泛评估表明,与传统企业基线和现有最优方法相比,SherAgent分别将端到端调查成功率提高了31.1%和63.7%。此外,它运行效率显著,每次调查的API成本低于0.10美元,耗时不到4分钟。最后,用户研究证实SherAgent提供准确清晰的见解,显著减少安全专家的分析负担。
英文摘要
Provenance-based attack investigation enables viable automation by standardizing data and query logic; however, it is critically hindered in practice by dependency explosions and fragmented causal chains in the wild. Towards designing a robust and automated investigation tool, we collaborated with the SOC of a major Internet corporation serving billions of users. By engaging in real-world incident response, we are able to evaluate and refine their existing LLM-based investigation workflows, which processes tens of thousands of raw alerts daily, leaving thousands for manual triage, to find out the root causes of investigation failures and major challenges in their existing tools. Motivated by these findings, we propose SherAgent, an LLM-empowered automated investigation system. Operating on an iterative ``query-filter'' backtracking paradigm over provenance graphs, SherAgent leverages the semantic reasoning capabilities of LLMs to process unstructured data, such as investigation context and threat intelligence. To overcome fragmented causal chains caused by missing events, the system dynamically calibrates query conditions to broaden the search scope. Concurrently, it performs precision result filtering and strategic nodes selection for subsequent exploration, thereby mitigating dependency explosions. Extensive evaluations in the wild demonstrate that SherAgent improves the end-to-end investigation success rate by 31.1% and 63.7% compared to both legacy enterprise baselines and SOTA approaches, respectively. Furthermore, it operates with remarkable efficiency, incurring under $0.10 in API costs and requiring less than 4 minutes per investigation. Finally, our user study confirms that SherAgent provides accurate and clear insights, significantly reducing the analytical overhead for security experts.