arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27311cs.CRcs.AI

多视图融合用于加密C2检测:泄漏受控的评估陷阱测量研究

Multi-View Fusion for Encrypted C2 Detection: A Leakage-Controlled Measurement Study of Evaluation Pitfalls

  • Academy of Cryptography Techniques(密码技术学院)
  • Ho Chi Minh City University of Technology (HCMUT)(胡志明市理工大学)
  • Vietnam National University Ho Chi Minh City(越南胡志明市国家大学)

机构由 AI 辅助整理,请以论文原文为准。

Hoang-Huy Nguyen-Huu, Van-Tri Phan, Khuong Nguyen-An

AI总结:

本研究通过泄漏受控的测量揭示加密C2检测评估中的三大陷阱:预处理泄漏、标签依赖及真实标签缺口,并指出评估设计本身才是核心结果。

AI中文摘要:

命令与控制(C2)流量日益隐藏于TLS之中,因此防御者现在将机器学习应用于流量元数据。许多研究假设结合两种元数据视图,即流统计和TLS握手指纹,能够同时提高准确性和鲁棒性。我们在来自62个真实Cobalt Strike捕获的17,577条TLS流上检验了这一假设。我们的评估消除了导致报告分数过于乐观的数据泄漏。我们报告了三个比融合结果本身更重要的发现。首先,一个不正确的预处理步骤使F1分数提高了0.28。该步骤在整个数据集上计算频率编码,而不是在每个交叉验证折内计算。这一提升比我们测量的任何真实效应大约大十倍。其次,标签和行为特征都依赖于目标地址。因此,17,577条流仅形成2,132个独立组,且55.1%的正样本率(看似平衡)降至4.2%。因此,类别平衡只是我们分析数据方式的结果,具体而言是我们计数流还是端点,而非任务的真实特征。第三,62个捕获中有20个(32%)没有到任何已知C2地址的TLS流,因此它们仅包含良性样本。我们直接检查了这些捕获,并确认这是真实标签中的缺口,而非标注错误。在此背景下,融合在F1上仅比最佳单视图高出0.022。当攻击者同时伪造两个特征面时,每个模型的表现都差于始终预测正类的简单基线(F1 = 0.711)。对于加密C2检测,评估设计并非初步步骤,它本身就是主要结果。

英文摘要:

Command-and-control (C2) traffic increasingly hides within TLS, so defenders now apply machine learning to traffic metadata. Many studies assume that combining two metadata views, namely flow statistics and TLS handshake fingerprints, improves both accuracy and robustness. We tested this assumption on 17,577 TLS flows from 62 real Cobalt Strike captures. Our evaluation removes the data leakage that leads to overly optimistic reported scores. We report three findings that matter more than the fusion result itself. First, an incorrect preprocessing step increases the F1 score by 0.28. This step computes the frequency encoding across the entire dataset rather than within each cross-validation fold. The increase is about ten times larger than any real effect we measured. Second, both the labels and the behavioral features depend on the destination address. Because of this, the 17,577 flows form only 2,132 independent groups, and the positive rate of 55.1\%, which looks balanced, drops to 4.2\%. Therefore, class balance is just a result of how we analyze the data, specifically whether we count flows or endpoints, and not a real feature of the task. Third, 20 of the 62 captures (32\%) have no TLS flows to any known C2 address, so they contain only benign samples. We checked these captures directly and confirmed that this is a gap in the ground truth, not a labeling error. In this context, fusion beats the best single view by only 0.022 in F1. When an attacker forges both feature surfaces simultaneously, every model performs worse than a simple baseline that always predicts positive (F1 = 0.711). For encrypted C2 detection, the evaluation design is not a preliminary step. It \emph{is} the main result.

补充信息

↑