发表机构
Department of Computer Science, San Jose State University; Faculty of Information Technology, Czech Technical University in Prague(圣何塞州立大学计算机科学系; 布拉格捷克理工大学信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于图像度量学习的恶意软件神经指纹方法,通过嵌入空间最近邻搜索实现零样本家族识别,在未见家族上取得73.1%检索@1和90.5%开放集AUROC。
AI 中文摘要
识别新观察到的恶意软件样本所属家族是威胁情报中的核心任务,然而传统分类器每当出现新家族时都必须重新训练。本章开发了一种基于图像的度量学习方法,该方法学习从恶意软件即图像表示中提取判别性神经指纹——固定长度的嵌入向量,从而可以通过嵌入空间中的最近邻搜索来确定家族归属。这种公式的核心优势在于零样本能力:因为学习到的嵌入诱导的是相似性度量而非固定的类别边界,训练期间从未见过的家族可以通过与图库比较而被识别,无需重新训练。我们直接证明了这一点:在MalNet-Images-Tiny和MalImg组合(453个家族,96,769张图像)上训练编码器,并在一个无家族重叠的留出17家族灰度数据集上对其进行零样本评估。使用带有多个代理锚点损失的轻量级CNN,该模型在编码器从未见过的家族上达到了73.1%的检索@1和90.5%的开放集AUROC。我们在同域设置中将我们的嵌入方法与两种传统范式进行基准测试,其中所有三种方法在将恶意软件分类到家族方面都具有竞争力。我们进一步表明,学习到的嵌入可以跨数据集迁移。与分类器不同,我们的嵌入方法还产生可解释的相似性得分,并通过Facebook AI相似性搜索(FAISS)扩展到大型图库。最后,我们使用检索@k、聚类纯度、轮廓系数、分离比、少样本准确率和开放集检测指标对学习到的嵌入空间进行了全面评估,并进行了图像扰动下的鲁棒性分析。
英文摘要
Identifying the family of a newly observed malware sample is a core task in threat intelligence, yet conventional classifiers must be retrained whenever a new family appears. This chapter develops an image-based metric learning approach that instead learns to extract discriminative neural fingerprints--fixed-length embeddings--from malware-as-image representations, so that family membership can be determined by nearest-neighbor search in the embedding space. The central advantage of this formulation is zero-shot capability: because the learned embedding induces a similarity metric rather than a fixed set of class boundaries, families that were never seen during training can be recognized by comparison against a gallery, with no retraining. We demonstrate this directly by training an encoder on MalNet-Images-Tiny and MalImg combined (453 families, 96,769 images) and evaluate it zero-shot on a held-out 17-family grayscale dataset with no family overlap. Using a lightweight CNN with multi-proxy anchor loss, this model attains 73.1% retrieval@1 and 90.5% open-set AUROC on families the encoder has never seen. We benchmark our embedding approach against two conventional paradigms in a same-domain setting, where all three are competitive at classifying malware into families. We further show that the learned embeddings transfer across datasets. Unlike classifiers, our embedding approach also yields interpretable similarity scores and scales to large galleries via Facebook AI Similarity Search (FAISS). Finally, we provide a comprehensive evaluation of the learned embedding space using retrieval@k, cluster purity, silhouette score, separation ratio, few-shot accuracy, and open-set detection metrics, along with robustness analysis under image perturbations.
CommentsTo appear as a chapter in the book "Artificial Intelligence for Cyber Defense in Emerging Threats", to be published by Springer by early 2027