发表机构
University of Kent(肯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究恶意软件分析中,小型语言模型编排集成能否超越单个大语言模型。通过测试多种模型建立基线,设计评估四种编排架构,其中混合系统表现出色,总体准确率超专业和前沿基线,证明基于证据的编排可提升SLM性能。
AI 中文摘要
恶意软件分析需要快速解读包含文件系统、网络和进程行为的复杂引爆报告。虽然大语言模型在技术工件解读方面能力令人印象深刻,但封闭权重前沿模型的不透明性和不断上升的API成本促使人们探索开放权重的替代方案。然而,许多开放权重模型很大,需要大量计算资源且托管成本高昂,资源受限的部署难以企及。本文研究小型语言模型(SLM)的编排集成是否能在关于恶意软件引爆报告的结构化问题上匹配或超越单个大语言模型的性能。通过在Meta的CyberSecEval恶意软件分析基准上测试11个开放权重SLM、3个网络安全预训练模型和6个前沿大语言模型建立基线。然后设计并评估了四种编排架构:多智能体管道、对抗性辩论框架、分层咨询系统和混合架构。混合系统(Qwen3 - 4B与Foundation - Sec - 8B)总体准确率达到35.30%,超过最强的网络安全专业基线(22.54%)和最强的无基础前沿基线(34.77%);在相同证据管道下,有基础的Gemini仍然是最强配置,准确率为38.22%。这些发现表明,基于证据的编排可以显著提高协作式SLM在支持恶意软件引爆报告解读方面的性能。
英文摘要
Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours. While large language models (LLMs) demonstrate impressive capabilities for technical artifact interpretation, the opacity and escalating API costs of closed-weight frontier models motivate exploration of open-weight alternatives. However, many open-weight models are large, demanding significant compute resources and incurring non-trivial hosting costs that place them beyond reach for resource-constrained deployments. This paper investigates whether orchestrated ensembles of small language models (SLMs) can match or exceed single LLM performance on structured questions about malware detonation reports. We established baselines by testing eleven open-weight SLMs, three cyber security pre-trained models, and six frontier LLMs on Meta's CyberSecEval Malware Analysis benchmark. We then designed and evaluated four orchestration architectures: (i) a multi-agent pipeline that decomposes analysis into structured evidence-collection and reasoning stages, (ii) an adversarial debate framework in which two agents iteratively critique each other's reasoning, (iii) a hierarchical consultation system that pairs a general-purpose SLM with a cyber-specialised expert model, and (iv) a hybrid architecture that combines evidence-grounded pipelines with adversarial debate reasoning. The hybrid system (Qwen3-4B with Foundation-Sec-8B) achieved 35.30% overall accuracy, exceeding the strongest cyber-specialised baseline (22.54%) and the strongest ungrounded frontier baseline (34.77%); when given the same evidence pipeline, grounded Gemini remained the strongest configuration at 38.22%. These findings show that evidence-grounded orchestration can substantially improve the performance of collaborative SLMs for supporting interpretation of malware detonation reports.
CommentsTo appear in Proceedings of the 29th International Symposium on Research in Attacks, Intrusions, and Defenses (RAID)