arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ANT:用于智能体行为审计的多粒度网络流量数据集与基准

ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing

Fan Li, Xiangyu Gao, Zixuan Liu, Tong Li, Chuanpu Fu, Ziqiang Wang, Ke Xu

arXiv 2610.06514首次发表:更新:

发表机构

Tsinghua University; Zhongguancun Laboratory; Renmin University of China; Nanyang Technological University(清华大学; 中关村实验室; 中国人民大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ANT数据集,包含多粒度标注的网络流量,用于审计LLM智能体行为,并建立基准评估现有方法,发现其在风险识别和场景区分上存在局限。

AI 中文摘要

大型语言模型(LLM)智能体的日益普及,使得网络管理员和安全团队需要在组织网络内审计智能体行为,同时避免检查用户的私人内容。网络流量提供了一种可观察的证据来源,但它在多大程度上揭示智能体任务和操作仍不清楚。现有的流量数据集缺乏评估该问题所需的联合任务和阶段标注。我们引入了ANT(智能体网络流量),这是一个在风险、场景和行为原语粒度上提供智能体行为信息以及网络流量的数据集。ANT包含20个任务和5个场景中的3,114个执行片段,由276,417个双向流和40,049个行为原语片段组成,这些片段被组织成47个宏组。我们建立了一个基准,用于智能体风险识别、场景识别和行为原语分类,使用了13个代表性的流量分析基线。结果表明,现有方法能恢复有用但不均衡的行为信号。当恶意工作流类似于良性任务时,它们难以识别风险,也难以区分具有相似流量模式的场景。对于频繁出现的宏组和具有独特流量模式的组,原语分类比对于罕见或语义相似的组更可靠。ANT为从网络流量中开发更精确的智能体行为审计和取证分析提供了共同基础。我们的数据和代码可在以下网址获取:此https URL。

英文摘要

The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic offers an observable source of evidence, but how much it reveals about agent tasks and operations remains unclear. Existing traffic datasets lack the joint task and stage annotations needed to evaluate this question. We introduce ANT (Agent Network Traffic), a dataset providing agent behavior information at risk, scenario, and behavior primitive granularities alongside network traffic. ANT contains 3,114 execution episodes across 20 tasks and five scenarios, comprising 276,417 bidirectional flows and 40,049 behavior primitive segments organized into 47 macro groups. We establish a benchmark for agent risk identification, scenario recognition, and behavior primitive classification using 13 representative traffic analysis baselines. The results show that existing methods recover useful but uneven behavioral signals. They struggle to identify risk when malicious workflows resemble benign tasks and to distinguish scenarios with similar traffic patterns. Primitive classification is more reliable for frequent macro groups and those with distinctive traffic patterns than for rare or semantically similar groups. ANT provides a common basis for developing more precise auditing and forensic analysis of agent behavior from network traffic. Our data and code are available at https://anonymous.4open.science/r/ant-main-suite-7BC0/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑