发表机构
Ontario Tech University; Western University(安大略理工大学; 韦仕敦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种LLM辅助的AutoML框架,通过LLM生成有界策略,实现物联网入侵检测的自动数据平衡、特征工程与算法选择,在低预算下提升F1分数并显著减少优化时间。
AI 中文摘要
物联网(IoT)系统正日益广泛地部署于智能家居、交通运输、能源系统及关键基础设施中。这种广泛的连接性提升了服务的智能化水平,但也扩大了物联网网络的攻击面。基于机器学习(ML)的入侵检测系统(IDS)被广泛用于识别恶意网络威胁并保护物联网系统,然而,开发有效的基于ML的IDS模型通常需要人类专业知识,并在许多流程(包括数据预处理、特征选择、模型选择和超参数调优)中反复进行手动决策。自动化机器学习(AutoML)通过使用优化技术自动化ML流水线的步骤来减轻这一负担,但传统的AutoML方法可能消耗大量的优化时间,因为它们探索了广泛的候选模型族和巨大的超参数空间。本文提出了一种面向物联网入侵检测的LLM辅助AutoML框架。该框架利用LLM作为策略生成器,将数据集概况转换为有界且经过验证的AutoML策略,用于自动数据平衡、自动特征工程以及组合算法选择与超参数优化(CASH)。在相同的10次试验预算下,所提出的LLM辅助策略在两个数据集上均实现了比使用树结构Parzen估计器(TPE)的传统AutoML更高的加权测试F1分数,在CICIDS2017上达到99.680%,在IoTID20上达到99.186%。相对于更广泛的30次试验的传统AutoML-TPE基线,所提出的10次试验方法分别将优化器时间减少了63.7%和49.9%,同时实现了略高的F1分数。这些结果表明,有界的LLM策略可以提高低预算AutoML搜索的质量,同时相对于更大的传统搜索预算保持明显的效率优势。
英文摘要
Internet of Things (IoT) systems are increasingly deployed in smart homes, transportation, energy systems, and critical infrastructure. This broad connectivity improves service intelligence, but also enlarges the attack surface of IoT networks. Machine Learning (ML)-based Intrusion Detection Systems (IDSs) are widely used to identify malicious network threats and protect IoT systems, but developing effective ML-based IDS models often requires human expertise and repeated manual decisions on many procedures, including data pre-processing, feature selection, model selection, and hyperparameter tuning. Automated Machine Learning (AutoML) reduces this burden by automating steps of the ML pipeline using optimization techniques, but conventional AutoML methods can consume substantial optimization time because they explore broad candidate model families and large hyperparameter spaces. This paper proposes a Large Language Model (LLM)-assisted AutoML framework for IoT intrusion detection. The proposed framework uses an LLM as a policy generator that converts dataset profiles into bounded and validated AutoML policies for automated data balancing, automated feature engineering, and Combined Algorithm Selection and Hyperparameter Optimization (CASH). Under an equal 10-trial budget, the proposed LLM-assisted policy achieves higher weighted test F1-score than traditional AutoML using the Tree-structured Parzen Estimator (TPE) on both datasets, reaching 99.680% on CICIDS2017 and 99.186% on IoTID20. Relative to the broader 30-trial Traditional AutoML-TPE baseline, the 10-trial proposed method reduces optimizer time by 63.7% and 49.9%, respectively, while achieving slightly higher F1-score. These results show that a bounded LLM policy can improve the quality of a low-budget AutoML search while retaining a clear efficiency advantage relative to a larger conventional search budget.
CommentsSubmitted to an IEEE Journal. Code will be released at: https://github.com/LiYangHart/LLM-Assisted-AutoML-For-Intrusion-Detection