arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05369cs.CRcs.AI

AutoDP-LLM:利用大型语言模型自动化入侵检测系统的数据预处理

AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models

Bao-Phong Nguyen, Gia-Khanh Pham, Thai-Duong Do, Mai Xuan Trang, Minh-Tuan Le, Xuan-Nam Tran, Huan Vu, Tien-Cuong Nguyen, Vu-Duc Ngo, Thien Van Luong

首次发表
浏览论文内容

中文总结 AI 辅助

AutoDP-LLM利用大型语言模型自动生成和验证入侵检测的数据预处理管道,在UNSW-NB15和NSL-KDD基准上实现竞争性检测性能,减少人工试错与计算开销。

中文摘要 AI 辅助

现代网络攻击的复杂性和规模日益增加,这要求入侵检测系统(IDS)具备智能且计算高效的能力。然而,传统上设计有效的数据预处理管道需要大量的试错工作,并反复评估不同的配置方案。对于大规模、高维的网络流量数据,这一过程可能造成显著的计算负担。在本工作中,我们提出了AutoDP-LLM,一个旨在减少人工管道开发和计算开销的自动化预处理框架。具体而言,AutoDP-LLM利用大型语言模型(LLMs)自主生成并验证可执行的数据预处理管道。该框架将确定性的主机端规划与基于LLM的专家代理相结合,以制定数据处理策略、合成可执行代码,并利用语义推理和训练得出的统计证据自适应地确定保留的特征集,而无需预定义特征预算。聚焦于多分类入侵检测,我们在UNSW-NB15和NSL-KDD基准数据集上,使用多个下游分类器对AutoDP-LLM进行了评估。与传统特征选择方法的对比实验表明,AutoDP-LLM在自动化生成紧凑且可执行的预处理管道的同时,实现了具有竞争力的检测性能。组件级消融实验进一步证明了语义和统计特征缩减组件的互补贡献。候选管道的重复生成、验证、执行和评估过程适合并行执行,这凸显了可扩展计算环境(包括高性能计算(HPC)系统)在支持自动化IDS管道开发方面的潜力。

英文摘要

The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems (IDS). However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort and repeated evaluation of alternative configurations. For large, high-dimensional network traffic data, this process can create a significant computational burden. In this work, we propose AutoDP-LLM, an automated pre-processing framework designed to reduce manual pipeline development and computational overhead. Specifically, AutoDP-LLM leverages Large Language Models (LLMs) to autonomously generate and validate executable data pre-processing pipelines. The framework combines deterministic host-side planning with LLM-based specialist agents to formulate data-processing strategies, synthesize executable code, and adaptively determine retained feature sets using semantic reasoning and training-derived statistical evidence, without requiring a predefined feature budget. Focusing on multiclass intrusion detection, we evaluate AutoDP-LLM on the UNSW-NB15 and NSL-KDD benchmark datasets using multiple downstream classifiers. Comparative experiments against conventional feature-selection methods show that AutoDP-LLM achieves competitive detection performance while automating the generation of compact and executable pre-processing pipelines. Component-level ablation experiments further demonstrate the complementary contributions of the semantic and statistical feature-reduction components. The repeated generation, validation, execution, and assessment of candidate pipelines are amenable to parallel execution, highlighting the potential of scalable computing environments, including high-performance computing (HPC) systems, to support automated IDS pipeline development.

发表机构

  • National Economics University(国民经济大学)
  • Phenikaa University(Phenikaa 大学)
  • Posts and Telecommunications Institute of Technology(邮电技术学院)
  • Le Quy Don Technical University(黎贵惇技术大学)
  • VNPT Group(越南邮政电信集团)
  • MobiFone Corp.(越南移动电信公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑