AI 中文总结
本文提出首个基于多智能体大语言模型的主动CTI获取系统DarkBot,其在受控与真实世界实验中均优于基线,可从地下论坛高效获取CTI且安全合规。
AI 中文摘要
从地下论坛获取的网络威胁情报(CTI)传统上依赖被动监控。然而,随着用户对大规模数据采集的意识增强,开放论坛中的有价值情报日益稀缺,往往转而迁移至私人或更难触及的空间,使得被动方法不再适用。基于可通过主动获取获取相关信息的直觉,本文提出DarkBot,据我们所知,这是首个用于地下论坛主动CTI获取的多智能体大语言模型(LLM)系统。DarkBot将交互任务分解为11个专门智能体,这些智能体分为三个功能模块:用于相关性和安全过滤的参与门控模块、由MITRE ATT&CK战术驱动的上下文感知问题生成模块,以及用于更好匹配真实论坛用户语言风格的语言风格适配模块。在针对100个CrimeBB对话的受控评估中,该系统仅通过观察每次交互开始时的初始帖子,就恢复了原始讨论中72.8%的经验证的MITRE ATT&CK技术,且始终优于单一基线系统。所提出的分层安全设计在流程层面遏制了所有注入的越狱尝试。真实世界实验进一步支持了这些结果:在一项前瞻性匹配部署中,分配给DarkBot的线程在7天内平均比对照组多积累3.85个CTI实体;在104个实时论坛对话中,该系统获取了与CTI相关的披露信息,且未观察到账户封禁、版主干预或对自动化参与的明确指责。
英文摘要
Cyber threat intelligence from underground forums has traditionally relied on passive monitoring. However, as users have become more aware of large-scale data collection, valuable intelligence has become increasingly rare in open forums, often migrating instead to private or harder-to-reach spaces, making passive approaches inadequate. Building on the intuition that relevant information can be obtained through active elicitation, this paper presents DarkBot, to the best of our knowledge, the first multi-agent LLM-based system for active CTI elicitation in underground forums. DarkBot decomposes the interaction task across eleven specialized agents organized into three functional blocks: engagement gating for relevance and safety filtering, context-aware question generation driven by MITRE ATT&CK tactics, and linguistic style adaptation to better align with real forum users. In a controlled evaluation across 100 CrimeBB conversations, the system recovered 72.8% of the validated MITRE ATT&CK techniques present in the original discussions by observing only the initial post at the start of each interaction, and it consistently outperformed a monolithic baseline. The proposed layered safety design contained all injected jailbreak attempts at the pipeline level. These results were further supported by real-world experiments: in a prospective matched deployment, threads assigned to DarkBot accumulated an average of 3.85 more CTI entities than their controls over seven days, and across 104 live forum conversations, the system elicited CTI-relevant disclosures without observed account suspensions, moderator interventions, or explicit accusations of automated participation.
Comments23 pages, 19 figures