用于恶意软件检测的控制流图神经网络的时间泛化与解释稳定性
Temporal Generalization and Explanation Stability of Control Flow Graph Neural Networks for Malware Detection
- Rajshahi University of Engineering & Technology(拉杰沙希工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过严格时间划分评估控制流图神经网络在恶意软件检测中的泛化能力,发现消息传递算子选择显著影响鲁棒性,且传统基准会误导模型选择,并据此提出无需搜索的架构。
AI中文摘要:
恶意软件检测是网络安全中的一项关键任务,基于控制流图的图神经网络已在该任务中展现出有前景的结果。然而,检测器通常在单一时间段内收集的语料库的随机划分上进行评估,这无法展示模型对后续样本的泛化能力。本研究通过严格的时间划分来解决这一局限性:每个模型在一个时间段上训练,并在稍后的时间段上进行一次评分。从1,989个Windows可移植可执行文件中静态提取了两个控制流图语料库,每个节点携带37个特征:其中459个图来自2024-2025年用于训练,223个图来自2026年用于评估。在较早的语料库上训练了十二个变体和一个扁平特征对照。消息传递算子的选择显著改变了应对分布偏移的鲁棒性,且每个在修正后仍存活的成对差距都将聚合架构与基于学习到的注意力读出架构区分开来。排名也发生了反转:扁平对照(仅看到节点特征而无拓扑结构)是分布内最佳模型,但在跨边界时属于最差模型之一,因此传统基准会拒绝消息传递。重新校准和集成都不能替代算子选择。归因没有偏移,但解释有效性是架构特定的,且在较晚语料库上最准确的算子最难解释。从该发现推导出的架构无需搜索即可匹配最佳搜索算子。偏移同时影响恶意软件和良性类别,因此这些结果是关于对分布偏移的鲁棒性,而非恶意软件演化。
英文摘要:
Malware detection is a critical task in cybersecurity, and graph neural networks over control flow graphs have shown promising results for it. However, detectors are usually evaluated on a random split of a corpus collected over a single period, which cannot show how well a model generalizes to later samples. This study addresses that limitation with a strict temporal split: every model is trained on one period and scored once on a later one. Two corpora of control flow graphs, each node carrying 37 features, were extracted statically from 1,989 Windows portable executables: 459 graphs from 2024-2025 for training and 223 from 2026 for evaluation. Twelve variants and a flat-feature control were trained on the earlier corpus. The choice of message-passing operator changes robustness to the shift significantly, and every pairwise gap that survives correction separates an aggregating architecture from one built around a learned attentional readout. The ranking also reverses: the flat control, which sees node features but no topology, is the best in-distribution model and among the worst across the boundary, so a conventional benchmark would have rejected message passing. Neither recalibration nor ensembling substitutes for the operator choice. Attributions do not shift, but explanation validity is architecture-specific, and the most accurate operator on the later corpus is the hardest to explain. An architecture derived from the finding matches the best searched operator without search. The shift affects both malware and benign classes alike, so these are results about robustness to distribution shift, not malware evolution.