发表机构
Purdue University; University of Texas at El Paso; Virginia Tech(普渡大学; 德克萨斯大学埃尔帕索分校; 弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RuleAutoPilot 是一个端到端智能体框架,直接从恶意软件网络流量生成可部署的 Suricata 规则,通过良性流量指纹识别降低噪声,并利用执行验证自动修复规则,显著提升规则质量并降低 LLM 成本。
AI 中文摘要
基于规则的入侵检测系统(IDS),如 Suricata,是网络安全的核心,然而制定有效的检测规则需要深厚的专家知识,且无法跟上新兴威胁的步伐。现有的基于大语言模型(LLM)的方法可以减少分析人员的工作量,但它们要么依赖于在底层流量工件已经存在之后才生成的精心策划的威胁情报,要么需要昂贵的 LLM 使用且缺乏足够的质量控制。我们提出了 RuleAutoPilot,一个端到端的智能体框架,直接从恶意软件网络流量生成可部署的 Suricata 规则,无需任何先验威胁情报。一个关键挑战是噪声:网络流量捕获通常包含少量与安全相关的流量,并混合有大量背景流量,这降低了 LLM 推理质量并增加了成本。RuleAutoPilot 通过一个良性流量指纹识别阶段来解决这一挑战,该阶段在 LLM 处理之前移除已知的良性背景流。未通过语法检查、未在源流量上触发或在良性语料库上产生误报的规则,将使用结构化反馈自动修复。在 1,296 个恶意软件 PCAP 上,基于执行的验证将规则质量(F1)从 0.443 提高到 0.539。在一个分层的 200-PCAP 子集上,使用开放权重模型 gpt-oss-120b 的 RuleAutoPilot 达到了接近前沿的质量,F1 为 0.524,而 Claude Code 下的 Claude Opus 5 为 0.623,且计费令牌成本降低了 52 倍。仅将骨干模型更换为 Claude Opus 5,RuleAutoPilot 便全面超越 Claude Code,F1 为 0.656 对比 0.623,且令牌使用量减少了 40 倍。更强的骨干模型提高了 RuleAutoPilot 自身的上限,但在相同的骨干模型下,我们的脚手架仍然优于 Claude Code 的,表明脚手架独立于骨干模型做出贡献。
英文摘要
Rule-based Intrusion Detection Systems (IDS) such as Suricata are central to network security, yet crafting effective detection rules demands deep expert knowledge and cannot keep pace with emerging threats. Existing LLM-based approaches can reduce analyst effort, but they either rely on curated threat intelligence that is produced only after the underlying traffic artifacts already exist, or they require costly LLM use without sufficient quality control. We present RuleAutoPilot, an end-to-end agentic framework that generates deployable Suricata rules directly from malware network traffic, with no prior threat intelligence required. A key challenge is noise: network traffic captures often contain a small amount of security-relevant traffic mixed with large volumes of background traffic, which reduces LLM reasoning quality and increases cost. RuleAutoPilot addresses this challenge with a Benign Traffic Fingerprinting stage that removes known benign background flows before LLM processing. Rules that fail syntax checks, do not trigger on the source traffic, or generate false positives on a benign corpus are automatically repaired using structured feedback. Across 1,296 malware PCAPs, execution-grounded verification raises rule quality (F1) from 0.443 to 0.539. On a stratified 200-PCAP subset, RuleAutoPilot on the open-weight gpt-oss-120b reaches near-frontier quality, 0.524 F1 against Claude Opus 5 under Claude Code's 0.623, at 52x lower billed-token cost. Swapping only the backbone to Claude Opus 5, RuleAutoPilot surpasses Claude Code outright, 0.656 F1 against 0.623, at 40x fewer tokens. A stronger backbone raises RuleAutoPilot's own ceiling, but at the same backbone, our scaffold still outperforms Claude Code's, showing the scaffold contributes independently of the backbone.