arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型在威胁等级判定中的基准测试

Benchmarking LLMs for Threat Level Determination

Han Wang, Murathan Kurfalı, Alfonso Iacovazzi

arXiv 2609.07582首次发表:更新:

发表机构

RISE Research Institutes of Sweden(瑞典RISE研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究构建MISP OSINT数据集,设计提示词并微调八种大语言模型,发现零样本性能弱而微调后F1达0.40-0.58,但仍需提升以用于实际威胁等级判定。

AI 中文摘要

大语言模型(LLMs)的快速发展为网络威胁情报管理开辟了新的机遇,但其在操作任务中的可靠性仍不明确。在本工作中,我们对大语言模型在威胁等级判定任务上进行了基准测试。首先,我们构建了一个源自公开可用的MISP OSINT源数据的精选数据集。接着,我们设计了一个定制的提示词,在零样本条件下系统性地比较了八种不同的大语言模型。最后,我们对每个模型应用了监督微调,并在基线与微调版本之间进行了比较分析。我们的结果表明,零样本模型性能较弱,正确分配威胁等级的能力有限。然而,微调后的模型展现出显著改进,根据基础架构的不同,F1分数达到0.40至0.58之间。尽管取得了这一进展,其性能对于实际部署而言仍然偏低,这凸显了在数据质量、模型适应和领域特定调优方面进行额外研究的必要性。

英文摘要

The fast progress of large language models (LLMs) opens new opportunities in the management of cyber threat intelligence, but their reliability for operational tasks remains unclear. In this work, we benchmark LLMs on the task of threat level determination. First, we construct a curated dataset derived from publicly available MISP OSINT feeds. Next, we design a tailored prompt to systematically compare eight different LLMs under zero-shot conditions. Finally, we apply supervised fine-tuning on each model and perform a comparative analysis between baseline and fine-tuned versions. Our results show that zero-shot models achieve weak performance, with limited ability to correctly assign threat levels. Fine-tuned models, however, demonstrate substantial improvements, reaching F1 scores between 0.40 and 0.58 depending on the base architecture. Despite this progress, the performance is still low for practical deployment, highlighting the need for additional research on data quality, model adaptation, and domain-specific tuning.

Journal ref2025 IEEE International Conference on Data Mining Workshops (ICDMW), Washington, DC, USA, 2025, pp. 1250-1257

DOI:10.1109/ICDMW69685.2025.00148

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑