arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过预算感知的流量标注在概念漂移下维持IoT设备识别

Maintaining IoT Device Identification under Concept Drift via Budget-Aware Traffic Labeling

Shayan Azizi, Norihiro Okui, Masataka Nakahara, Ayumu Kubota, Gustavo Batista, Hassan Habibi Gharakaheili

arXiv 2608.15465首次发表:更新:

发表机构

School of Electrical Engineering and Telecommunications, University of New South Wales; School of Computer Science and Engineering, University of New South Wales; KDDI Research(新南威尔士大学电气工程与电信学院; 新南威尔士大学计算机科学与工程学院; KDDI研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对IoT设备识别中概念漂移导致的分类性能下降问题,提出结合均匀流量采样与漂移检测确定标注量的策略,经380万条IPFIX流记录验证,可有效维持分类性能。

AI 中文摘要

从被动流量中识别IoT设备类型正越来越多地用于企业和ISP网络的安全管理。然而,由于设备行为不断演变,基于机器学习的分类器性能会在概念漂移下逐渐下降。因此,维持分类性能需要定期用新标注的部署流量进行重新训练。操作层面的挑战在于确定应标注多少以及哪些部署流量实例,以维持分类性能。我们表明,这两个决策应分开处理:仅基于漂移检测器选择的实例进行重新训练容易系统地忽略新兴行为空间的部分内容,而均匀采样的部署流量能捕捉更具代表性的行为变化;相反,漂移检测更适合确定应标注的部署流量数量。我们做出三项贡献:(1)对IoT流量进行了为期两年的纵向研究,刻画了行为演变如何在不同设备类别中显现,以及用新标注流量重新训练如何恢复分类性能;(2)开发了一种基于一致性的漂移检测器,可直接从原始流量特征中捕捉类别条件行为模型,并提供行为演变的特征级解释;(3)我们证明,根据观察到的行为演变调整流量标注率,结合均匀流量采样,能比检测器引导的样本选择更有效地维持分类器性能,且有助于管理流量标注工作量;我们进一步表明,该策略的性能与置信度引导的适应相当,同时提供特征级解释。我们的评估使用了从21种IoT类型、历时2年多收集的380万条IPFIX流记录。

英文摘要

Identification of IoT device types from passive traffic is increasingly used for security management in enterprise and ISP networks. However, the performance of machine learning-based classifiers gradually degrades under concept drift as device behavior evolves. Therefore, maintaining classification performance requires periodic retraining with newly labeled deployment traffic. The operational challenge is determining how much and which deployment traffic instances to label for maintaining classification performance. We show that these two decisions should be treated separately. While retraining solely on instances selected by a drift detector is prone to systematically overlooking parts of the emerging behavioral space, uniformly sampled deployment traffic captures more representative behavioral changes. Instead, drift detection is more effective at determining the amount of deployment traffic that should be labeled. We make three contributions. (1) We conduct a two-year longitudinal study of IoT traffic and characterize how behavioral evolution manifests across device classes and how retraining with newly labeled traffic restores classification performance. (2) We develop a conformity-based drift detector that captures class-conditional behavioral models directly from raw traffic features and provides feature-level explanations of behavioral evolution. (3) We demonstrate that adjusting the traffic labeling rate according to the observed behavioral evolution, combined with uniform traffic sampling, maintains classifier performance more effectively than detector-guided sample selection and is beneficial to managing the traffic labeling effort. We further show that this strategy performs comparably to confidence-guided adaptation while providing feature-level explanations. Our evaluation uses 3.8 million IPFIX flow records collected from 21 IoT types over more than 2 years.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑