发表机构
University of Technology Sydney; Yunnan University; University of New South Wales(悉尼科技大学; 云南大学; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对安卓恶意软件概念漂移,提出HYDRA框架,利用层次图对比学习对齐历史与目标数据分布,统一特征学习与域对齐,显著降低误报率并减少最多87.5%标注需求。
AI 中文摘要
概念漂移由安卓恶意软件的快速演变驱动,严重降低了机器学习检测器的性能。当前的适应策略往往是反应式的,仅在性能下降后才做出响应,并施加了显著的人工标注负担;或者它们是主动式的,但依赖于不稳定的对抗训练和不完整的、单层次的图表示。为克服这些局限,我们提出了HYDRA(混合漂移适应),一种主动式适应框架,从层次结构化数据中学习漂移不变表示。HYDRA首先使用混合图结构对应用程序进行建模,结合细粒度的控制流图(CFGs)和粗粒度的函数调用图(FCGs)以捕获全面的行为模式。随后,它引入了一种新颖的跨域对比学习目标,该目标对齐历史(源)数据和新(目标)数据的分布。通过为未标注的目标样本生成伪标签,我们的方法在单一、稳定的优化过程中,将语义相似的应用程序的表示拉近,无论其所属域如何。这种方法统一了特征学习和域对齐,无需复杂的对抗目标。在大型、按时间排序的恶意软件数据集上进行的大量实验表明,HYDRA实现了比最先进的基线显著更低的假阴性和假阳性率,同时所需的标注样本最多减少87.5%。因此,我们的工作为对抗安全应用中的概念漂移提供了一种稳健且高效的解决方案。
英文摘要
Concept drift, driven by the rapid evolution of Android malware, severely degrades the performance of machine learning detectors. Current adaptation strategies are often reactive, responding only after performance has dropped and imposing a significant manual annotation burden, or they are proactive but rely on unstable adversarial training and incomplete, single-level graph representations. To overcome these limitations, we propose HYDRA (Hybrid Drift Adaptation), a proactive adaptation framework that learns drift-invariant representations from hierarchically structured data. HYDRA first models applications using a hybrid graph structure, combining fine-grained Control Flow Graphs (CFGs) and coarse-grained Function Call Graphs (FCGs) to capture comprehensive behavioral patterns. It then introduces a novel cross-domain contrastive learning objective that aligns historical (source) and new (target) data distributions. By generating pseudo-labels for unlabeled target samples, our method pulls representations of semantically similar applications together, regardless of their domain, within a single, stable optimization process. This approach unifies feature learning and domain alignment, eliminating the need for complex adversarial objectives. Extensive experiments on large-scale, time-ordered malware datasets demonstrate that HYDRA achieves substantially lower False Negative and False Positive Rates than state-of-the-art baselines while requiring up to 87.5% fewer labeled samples. Our work thus offers a robust and efficient solution to combat concept drift in security applications.
CommentsAccepted at ACM CCS 2026. Author's version with full appendix. 17 pages