武器化地面真值:利用杀毒软件与基于学习的检测器之间的边界错位进行数据投毒攻击
Weaponizing Ground Truth: Data Poisoning Attacks by Exploiting Boundary Misalignment Between Antivirus Software and Learning-Based Detectors
- Nankai University(南开大学)
- Singapore Management University(新加坡管理大学)
- University of Liverpool(利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对ML恶意软件检测依赖AV标签的弱点,提出Bi-Iocane黑盒投毒框架,通过修改AV敏感字节翻转标签,以极低投毒预算使下游模型对92.08%目标误分类,且现有防御效果有限。
AI中文摘要:
基于机器学习(ML)的恶意软件检测器通常使用从杀毒软件(AV)引擎和聚合服务(如VirusTotal)获得的标签进行训练。这种做法假设AV生成的标签提供了可靠的监督。然而,微小的字节级修改可以大幅改变AV的判定,同时使下游ML检测器感知的表示基本保持不变,从而产生标签-特征不一致性,这可能污染训练数据集,并为基于ML的恶意软件检测创造投毒机会。我们提出了Bi-Iocane,一个黑盒投毒框架,利用恶意软件标签流水线对AV生成标签的依赖。Bi-Iocane识别AV敏感字节并修改它们以诱导标签变化。它在恶意软件中重写此类字节以获得良性标签(逃避导向的投毒),并将恶意软件关联的字节模式注入良性软件以获得恶意标签(诽谤导向的投毒)。这些投毒样本及其轻微修改的变体污染训练数据,并导致选定的目标被错误分类。我们使用13个AV引擎模拟AV聚合服务和8个ML检测器评估了Bi-Iocane。对于30个恶意软件和30个良性干净目标,Bi-Iocane结合AV特定操作生成恶意软件到良性和良性到恶意软件的投毒样本,其所有测试的基于AV的标签都被翻转。在这些投毒样本和变体用于下游训练后,生成的ML模型平均将92.08%的原始干净目标错误分类,每个目标仅需0.06%的投毒预算。同时,投毒模型在很大程度上保持了干净集性能,并且六种评估的投毒防御仅显示出有限的缓解效果。VirusTotal评估进一步证实了实际的诽谤风险,并揭示了现实世界AV到ML标签供应链中潜在的逃避风险。
英文摘要:
Machine-learning (ML)-based malware detectors are commonly trained using labels obtained from antivirus (AV) engines and aggregation services (e.g., VirusTotal). This practice assumes AV-generated labels provide reliable supervision. However, small byte-level modifications can substantially alter AV verdicts while leaving the representations perceived by downstream ML detectors largely unchanged, producing label-feature inconsistencies that can contaminate training datasets and create poisoning opportunities for ML-based malware detection. We present Bi-Iocane, a black-box poisoning framework that exploits the reliance of malware-labeling pipelines on AV-generated labels. Bi-Iocane identifies AV-sensitive bytes and modifies them to induce label changes. It rewrites such bytes in malware to obtain benign labels (evasion-oriented poisoning) and injects malware-associated byte patterns into benign software to obtain malicious labels (defamation-oriented poisoning). These poisoned samples and their lightly modified variants corrupt training data and cause selected targets to be misclassified. We evaluate Bi-Iocane with 13 AV engines simulating AV aggregation services and eight ML detectors. For 30 malware and 30 benign clean targets, Bi-Iocane combines AV-specific manipulations to generate malware-to-benign and benign-to-malware poisoned samples whose all tested AV-based labels are flipped. After these poisoned samples and variants are used for downstream training, the resulting ML models misclassify 92.08% of the original clean targets on average with only a 0.06\% poisoning budget per target. Meanwhile, the poisoned models largely preserve clean-set performance, and six evaluated poisoning defenses show only limited mitigation. VirusTotal evaluation further confirms practical defamation risk and reveals potential evasion risk in real-world AV-to-ML labeling supply chains.