AI 中文总结
本研究提出验证引导式规约合成框架,结合LLM与CEGIS从HTTP请求轨迹生成Suricata规则,在281个真实CVE及良性流量实验中,检测率达81.5%且误报率为0.0%。
AI 中文摘要
针对联网物联网设备的攻击持续增加,但将观测到的攻击流量转化为可部署的入侵检测系统(IDS)规则在很大程度上仍是手动过程。近期研究探索使用大语言模型(LLM)生成IDS规则,但现有方法通常需要观测流量之外的辅助信息,或生成规则后未针对良性流量验证其检测逻辑。本研究提出一种验证引导式规约合成框架,可直接从HTTP请求轨迹生成Suricata规则。该方法并非让LLM单步生成IDS规则,而是先由LLM识别易受攻击的参数并合成语义检测规约,随后通过反例引导的归纳合成(CEGIS)迭代优化这些规约,其中良性流量样本在合成与验证过程中充当反例。经验证的规约随后被确定性编译为Suricata规则。在281个真实世界的CVE及从真实物联网设备收集的良性流量上开展的实验表明,所提方法的检测率达81.5%,同时保持0.0%的误报率;一项消融研究还显示,基于CEGIS的验证可提升检测性能,同时维持低误报率。
英文摘要
Attacks against Internet-connected IoT devices continue to increase; however, transforming observed attack traffic into deployable intrusion detection system (IDS) rules remains largely a manual process. Recent studies have explored using large language models (LLMs) to generate IDS rules; nonetheless, existing approaches often require auxiliary information beyond observed traffic or generate rules without validating their detection logic against benign traffic. This study presents a verification-guided specification synthesis framework for generating Suricata rules directly from HTTP request traces. Instead of having an LLM generate IDS rules in a single step, an LLM first identifies a vulnerable parameter and synthesizes a semantic detection specification. These specifications are iteratively refined through counterexample-guided inductive synthesis (CEGIS), in which benign traffic samples serve as counterexamples during synthesis and verification. Verified specifications are then deterministically compiled into Suricata rules. Experiments on 281 real-world CVEs and benign traffic collected from real IoT devices show that the proposed method achieves a detection rate of 81.5% while maintaining a false positive rate of 0.0%. An ablation study also demonstrates that CEGIS-based verification improves detection performance while maintaining a low false positive rate.
CommentsAccepted at CIKM 2026