arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02348cs.CRcs.LG

用于二进制程序聚类的自监督表示:从实证研究到检索增强学习

Self-Supervised Representations for Binary Program Clustering: From Empirical Study to Retrieval-Augmented Learning

  • Faculty of Information Technology, Brno University of Technology(布尔诺理工大学信息技术学院)
  • Kempelen Institute of Intelligent Technologies (KInIT)(肯佩伦智能技术研究所(KInIT))
  • Faculty of Electrical Engineering and Information Technology, Slovak University of Technology(斯洛伐克理工大学电气与信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Martin Mocko, Daniela Chudá

AI总结:

本研究系统探究SSL和TRL在二进制程序聚类中的应用,提出VIME-R方法,在Ember和Bodmas数据集上实现同质性提升2.7%-5.8%,为自动化恶意软件分析提供新方向。

AI中文摘要:

恶意软件聚类是网络安全领域的关键任务,有助于发现威胁并分析不断演变的恶意软件家族。尽管自监督学习(SSL)和表格表示学习(TRL)在其他领域已取得突破,但它们在二进制程序聚类(对所有传入样本进行聚类,无论标签如何)中的应用仍基本未被探索。本研究首次对用于二进制程序聚类的SSL和TRL方法进行系统研究,分两个阶段在公开的Ember和Bodmas数据集上开展。第一阶段,我们通过采用有监督对生成将著名的基于视觉的SSL模型(BYOL、SimSiam、Barlow Twins、VICReg)适配到表格数据,确定了性能上限,发现BYOL和SimSiam的性能可与全监督模型媲美,而Barlow Twins和VICReg的性能则显著落后。第二阶段,我们评估纯无监督TRL方法与强基线(PCA、Autoencoder、UMAP)的表现,结果显示VIME成为二进制程序聚类的新标杆。基于这些发现,我们提出VIME-R,这是VIME的检索增强扩展,它用基于检索的增强替代随机边际分布损坏,以生成更具信息性的训练对。VIME-R的性能较VIME进一步提升,在两个数据集上的同质性指标提高了2.7%至5.8%。我们的结果表明,检索增强的表格表示学习是增强自动化恶意软件分析的有前景方向。代码将公开提供。

英文摘要:

Malware clustering is a critical task in cybersecurity that helps discover threats and analyze evolving malware families. While self-supervised learning (SSL) and tabular representation learning (TRL) have achieved breakthroughs in other domains, their application to binary program clustering (the task of clustering all incoming samples regardless of label) remains largely unexplored. This study presents the first systematic investigation of SSL and TRL methods for binary program clustering, conducted in two phases on the public Ember and Bodmas datasets. In Phase 1, we establish a performance ceiling by adapting prominent vision-based SSL models (BYOL, SimSiam, Barlow Twins, VICReg) for tabular data with supervised pair generation, finding that BYOL and SimSiam achieve performance comparable to fully supervised models, while Barlow Twins and VICReg significantly underperform. In Phase 2, we evaluate purely unsupervised TRL methods against strong baselines (PCA, Autoencoder, UMAP), demonstrating that VIME establishes a new state of the art for binary program clustering. Informed by these findings, we propose VIME-R, a retrieval-augmented extension of VIME that replaces random marginal-distribution corruption with retrieval-based augmentation to generate more informative training pairs. VIME-R further improves upon VIME, achieving 2.7\%-5.8\% higher Homogeneity on both datasets. Our results highlight retrieval-augmented tabular representation learning as a promising direction for enhancing automated malware analysis. Code will be made available.

↑