发表机构
Virginia Tech; UC Berkeley(弗吉尼亚理工大学; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NetAgent提出首个多任务智能体网络流量分析框架,通过知识增强规划、150+工具和三层记忆等设计,在9个基准上超越现有方法,显著提升泛化性与鲁棒性。
AI 中文摘要
网络流量分析是网络安全的基石,涵盖从入侵检测到加密流量分类等众多任务。现有方法要么训练任务特定的模型,但其泛化能力较差;要么依赖成本高昂的流量基础模型,而这些模型在分布偏移下仍然表现不佳。我们提出了NetAgent,这是首个用于多任务流量分析的智能体框架。通过精心设计的智能体循环,NetAgent支持复杂任务理解、即时分解与编排、动态重新规划以及长时程分析,且无需任务特定训练。它引入了五项关键设计:(1)知识增强的工作流规划,将攻击知识映射到流量特征,以弥合语义鸿沟;(2)全面的工具动作空间,包含从50多个已发表系统中提取的150多个经过验证的工具;(3)统一的代码执行空间,用于灵活的动作组合;(4)三层记忆结构,用于长期知识整合;(5)沙箱与运行时修复,确保可靠执行。在9个主要基准测试中,NetAgent在几乎所有任务上均优于所有基线(23个单任务和5个多任务),并且在未见流量分布上泛化能力显著更优(F1分数为90.04%,而最佳单任务和多任务基线分别为2.74%和3.04%),在现实背景偏移下也表现稳健(F1下降4.85个百分点,而最佳单任务和多任务基线分别下降74.88和74.80个百分点)。这些结果表明,现有方法所报告的许多成功源于对数据集特定模式的过拟合,在现实网络环境中性能急剧下降,而NetAgent的智能体设计则保持准确、泛化能力强且稳健。
英文摘要
Network traffic analysis is central to network security, spanning tasks from intrusion detection to encrypted traffic classification. Existing approaches either train task-specific models that generalize poorly or rely on costly traffic foundation models that still struggle under distribution shift. We present NetAgent, the first agentic framework for multi-task traffic analysis. Through a carefully designed agent loop, NetAgent supports complex task understanding, on-the-fly decomposition and orchestration, dynamic replanning, and long-horizon analysis, without task-specific training. It introduces five key designs: (1) knowledge-augmented workflow planning that maps attack knowledge to traffic features to bridge the semantic gap; (2) a comprehensive tool action space with 150+ verified tools extracted from 50+ published systems; (3) a unified code execution space for flexible action composition; (4) a three-tier memory for long-term knowledge consolidation; and (5) sandboxing and runtime repair for reliable execution. Across 9 major benchmarks, NetAgent outperforms all baselines (23 single-task and 5 multi-task) on nearly all tasks and generalizes substantially better to unseen traffic distribution (90.04% F1 vs. 2.74% and 3.04% for the best single-task and multi-task baselines) and under realistic background shift (4.85-point F1 drop vs. 74.88-point and 74.80-point drop for the best single-task and multi-task baselines). These results reveal that existing methods owe much of their reported success to overfitting dataset-specific patterns and degrade sharply in realistic network environments, while NetAgent's agentic design remains accurate, generalizable, and robust.
Comments23 pages,8 figures