以行为为中心的恶意软件分类与细粒度恶意逻辑定位
Behavior-Centric Malware Classification with Fine-Grained Malicious Logic Localization
- Gianforte School of Computing, Montana State University(蒙大拿州立大学吉安福特计算机学院)
- College of Computer Science and Technology, Shandong University(山东大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种以行为为中心的恶意软件分类框架,通过上下文敏感切片和行为图表示,结合Transformer与图神经网络,实现细粒度的恶意逻辑定位与准确分类。
AI中文摘要:
有效的恶意软件分析不仅需要了解程序是否具有恶意性,还需要了解其表现出哪些行为以及这些行为在代码中的起源位置。现有的基于机器学习的恶意软件检测器大多作为黑盒运行,对其决策所依据的恶意逻辑提供的洞察有限。本文解决了基本块级别的恶意行为定位与分类问题。我们提出了一种以行为为中心的分析框架,该框架将恶意软件样本分解为行为,并将这些行为系统地关联到其起源的代码区域。通过从与安全相关的系统API调用进行上下文敏感的后向切片,我们重建了控制依赖和数据依赖链,并将每个行为表示为相关基本块的结构化图。基于Transformer的模型捕获指令级语义,而图神经网络则对行为图中的结构依赖进行建模。所得到的表示被融合以实现准确且可解释的分类,基于注意力的归因识别负责恶意行为的代码区域。我们使用标准分类指标以及衡量手动标记的恶意行为检测率的行为覆盖率指标来评估我们的方法。我们的结果表明,所提出的框架在提供细粒度、行为感知的恶意逻辑定位的同时,实现了准确的恶意软件分类。
英文摘要:
Effective malware analysis requires understanding not only whether a program is malicious, but also which behaviors it exhibits and where those behaviors originate in the code. Existing machine-learning-based malware detectors largely operate as black boxes, providing limited insight into the malicious logic responsible for their decisions. This paper addresses malicious behavior localization and classification at the basic-block level. We propose a behavior-centric analysis framework that decomposes malware samples into behaviors and systematically links these behaviors to their originating code regions. Using context-sensitive backward slicing from security-relevant system API calls, we reconstruct control- and data-dependency chains and represent each behavior as a structured graph of related basic blocks. A Transformer-based model captures instruction-level semantics, while a Graph Neural Network models structural dependencies within behavior graphs. The resulting representations are fused to enable accurate and interpretable classification, with attention-based attribution identifying code regions responsible for malicious behaviors. We evaluate our approach using standard classification metrics and a behavior coverage metric that measures the detection of manually labeled malicious behaviors. Our results demonstrate that the proposed framework achieves accurate malware classification while providing fine-grained, behavior-aware localization of malicious logic.