arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27357cs.CR

SAGEGAN:基于高斯嵌入的生成对抗网络风格异常检测

SAGEGAN: Style-Based Anomaly Detection with Gaussian Embeddings using Generative Adversarial Networks

Thesath Wijayasiri, Kar Wai Fok, Vrizlynn L. L. Thing

首次发表
浏览论文内容

中文总结 AI 辅助

针对恶意软件快速演化问题,提出SAGEGAN框架,将可执行文件转为图像,用风格条件对抗重建建模良性结构,实现异常检测,在214家族语料上AUC达89.76%。

中文摘要 AI 辅助

恶意软件的演化速度超过了基于规则和签名驱动的检测流程。本文提出了SAGEGAN,一个仅使用良性样本训练的恶意软件异常检测框架,该框架将可移植可执行文件转换为紧凑的三通道图像,并通过风格条件对抗重建来建模良性结构。该表示结合了希尔伯特映射的字节值、良性参考的字节转换突变性以及相对于良性软件的熵偏差。模型将每个图像编码为与七阶段调制生成器对齐的逐层风格张量,而非单一潜在瓶颈。采用高斯风格先验、基于矩的先验对齐和潜在一致性来减少编码的良性风格与生成器采样流形之间的不匹配。为了解释性,一个确定性的编码器路径将每个可执行文件映射到固定的风格张量,从而实现可重复的逐层家族距离、梯度敏感性、主成分和类别行为分析。在自收集的包含214个家族恶意软件的可移植可执行文件语料库上,高斯变体实现了89.76%的接收者操作特征曲线下面积和88.19%的平衡准确率,而基因组风格变体分别达到88.03%和84.09%。在不重新拟合模型权重、良性参考统计或决策阈值的情况下,相同的检查点在DIKE、Microsoft BIG 2015和Lester恶意软件子集上进行了评估。结果表明,逐层风格建模既支持异常排名,也支持对恶意软件家族如何偏离良性流形的结构化事后分析。

英文摘要

Malware evolves faster than rule-based and signature-driven detection pipelines. This paper presents SAGEGAN, a benign-only trained malware anomaly detection framework that converts portable executable files into compact three-channel images and models benign structure through style-conditioned adversarial reconstruction. The representation combines Hilbert-mapped byte values, benign-referenced byte-transition surprise, and entropy deviation from benign software. The model encodes each image into a layer-wise style tensor aligned with a seven-stage modulated generator, rather than a single latent bottleneck. A Gaussian style prior, moment-based prior alignment, and latent consistency are used to reduce mismatch between encoded benign styles and the generator's sampled manifold. For interpretation, a deterministic encoder pathway maps each executable to a fixed style tensor, enabling repeatable layer-wise family distance, gradient sensitivity, principal component, and class-behaviour analyses. On a self-collected portable executable corpus containing malware from 214 families, the Gaussian variant achieves 89.76% area under the receiver operating characteristic curve and 88.19% balanced accuracy, while the genome-style variant reaches 88.03% and 84.09%, respectively. Without refitting model weights, benign reference statistics, or decision thresholds, the same checkpoints are evaluated on DIKE, Microsoft BIG 2015, and Lester malware subsets. The results suggest that layer-wise style modelling supports both anomaly ranking and structured post hoc analysis of how malware families depart from the benign manifold.

发表机构

  • Cybersecurity Strategic Technology Center, Singapore Technologies Engineering, Singapore(新加坡科技工程公司网络安全战略技术中心)

机构由 AI 辅助整理,请以论文原文为准。

↑