AI 中文总结
TagZilla是基于LLM的IoC自动标注平台,采用开放与闭世界分类法,在100份含1534个IoC的报告上F1值达0.94和0.93,可标注多行为者报告并生成IoC档案。
AI 中文摘要
网络威胁情报(CTI)报告常描述网络攻击涉及的恶意代码指示器(IoC),如IP地址、URL、文件哈希值及加密货币钱包。这些IoC通常以非结构化文本形式呈现于报告中,或在报告末尾列出但缺乏上下文,限制了其可用性。本文提出TagZilla平台,给定一份威胁报告,可自动分析其文本,为所描述的IoC标注上下文信息,包括该IoC所属的威胁组织、恶意软件家族,以及关联的滥用类型(如钓鱼、性勒索、命令与控制)。TagZilla采用一种新颖的基于大语言模型(LLM)的开放世界分类方法,为IoC分配所有者标签;同时采用闭世界分类法,为IoC分配29种滥用类型标签。我们在人工生成的包含100份威胁报告(含1534个指示器)的基准真值数据集上对TagZilla进行评估,其所有者标注的F1值达0.94,滥用类型标注的F1值达0.93。随后,我们将TagZilla应用于标注765份威胁报告,识别出15583个IoC,这些IoC属于637个恶意软件家族、113个威胁组织及162个其他实体。结果表明,TagZilla即便在描述多个行为者和恶意软件家族的报告中也能标注IoC,支持为这些实体生成IoC档案。
英文摘要
Cyber Threat Intelligence (CTI) reports often describe Indicators of Compromise (IoCs) such as IP addresses, URLs, file hashes, and cryptocurrency wallets involved in cyberattacks. Those IoCs are typically described in the unstructured report's text, or listed at the end of the report with little context, limiting their usefulness. This paper presents TagZilla, a platform that, given a threat report, automatically analyzes its text and tags the IoCs it describes with contextual information about the threat group and malware family that the IoC belongs to and the type of abuse associated with the IoC (e.g., phishing, sextortion, command-and-control). TagZilla provides a novel LLM-based approach to assign owner tags to IoCs using an open-world classification, and assigns 29 abuse type tags to IoCs using a closed-world classification. We evaluate TagZilla on a manually generated ground truth of 100 threat reports containing 1,534 indicators, where it achieves an F1 score of 0.94 for owner tagging and 0.93 for abuse type tagging. Then, we apply TagZilla to tag 765 threat reports, identifying 15,583 IoCs belonging to 637 malware families, 113 threat groups, and 162 other entities. The results show that TagZilla can tag IoCs even in reports describing multiple actors and malware families, enabling the generation of IoC profiles for those entities.