arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCRIPTIOC-BENCH:一个用于评估基于脚本的恶意软件中可操作威胁情报识别能力的LLM基准

SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs

Hanna Kim, Jian Cui, Minkyoo Song, Hwanjo Heo, Seungwon Shin, Kimin Lee, Xiaojing Liao

arXiv 2609.06149首次发表:更新:

发表机构

KAIST; University of Illinois Urbana-Champaign; ETRI(韩国科学技术院; 伊利诺伊大学厄巴纳-香槟分校; 韩国电子通信研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对脚本恶意软件中IOC静态恢复难题,提出包含634个样本的基准SCRIPTIOC-BENCH,评估多种LLM,最强模型F1仅65.4,并引入误报分类法及两种缓解措施以提升恢复性能。

AI 中文摘要

基于脚本的恶意软件仍然是一种普遍的攻击技术。这些脚本通常包含危害指标(IOCs),可提供可操作的威胁情报。然而,静态恢复此类指标具有挑战性,因为相关值可能在代码中分散或转换。尽管大型语言模型(LLMs)在安全分析中显示出潜力,但其从恶意脚本中恢复IOCs的能力仍未得到充分探索。我们提出了SCRIPTIOC-BENCH,一个用于衡量在真实世界恶意脚本上静态IOC提取能力的基准。该基准包含634个经人工验证的JavaScript、PowerShell和VBScript恶意软件样本,涵盖四种IOC类型(URL、域名、IP地址和文件系统工件)。我们进一步按恢复级别对真实IOC进行分层,区分直接暴露的指标与需要解码或重建的指标。利用此基准,我们评估了广泛的专有和开放权重LLMs,并表明在不执行的情况下恢复IOC在模型规模上仍然具有挑战性:最强的模型仅达到65.4的F1分数。为了描述恢复失败的特征,我们引入了一个误报分类法,并用它来比较所评估模型的错误概况。我们进一步在一个小型开放权重模型上研究了两种缓解措施,即确定性字符串实用程序和任务特定适配,发现它们提供了互补的恢复增益,提高了精确度,并将错误转向基于样本的不匹配。

英文摘要

Script-based malware remains a prevalent attack technique. These scripts often contain indicators of compromise (IOCs) that provide actionable threat intelligence. However, statically recovering such indicators is challenging, as relevant values may be dispersed or transformed within code. Although large language models (LLMs) have shown promise in security analysis, their ability to recover IOCs from malicious scripts remains underexplored. We present SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts. The benchmark comprises 634 manually verified JavaScript, PowerShell, and VBScript malware samples covering four IOC types (URLs, domains, IP addresses, and filesystem artifacts). We further stratify ground-truth IOCs by recovery level, distinguishing directly exposed indicators from those requiring decoding or reconstruction. Using this benchmark, we evaluate a broad range of proprietary and open-weight LLMs and show that IOC recovery without execution remains challenging across model scales: the strongest model reaches only 65.4 F1. To characterize how recovery fails, we introduce a false-positive taxonomy and use it to compare the error profiles of the evaluated models. We further study two mitigations on a small open-weight model, deterministic string utilities and task-specific adaptation, finding that they provide complementary recovery gains, raise precision, and shift errors toward sample-grounded mismatches.

Comments22 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑