发表机构
Nexcepta(Nexcepta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对通用LLM生成检测规则质量差的问题,提出领域自适应模型Sigma-Hunter,基于3,635条规则微调7B模型,在语义质量上超越通用基线,支持本地部署。
AI 中文摘要
检测工程师必须将威胁报告、取证观察和狩猎假设转化为精确、可测试的规则。通用大语言模型(LLM)可以起草此类规则,但常常生成无效的YAML、错误的日志源、不支持的字段或过于宽泛的检测逻辑。本文提出了Sigma-Hunter,一个面向分析师辅助的Sigma规则生成和威胁狩猎的领域自适应LLM。我们从3,635条经过验证的开源Sigma规则构建了一个指令微调数据集,扩展为7,663个问答和分析师推理示例。在扩展之前,每条源规则被分配到单一的训练、验证或测试分区,因此没有规则在分区之间泄漏。我们使用LoRA对7B的Mistral模型和Phi-4模型进行微调,并在语法、近似字段一致性以及检测逻辑、完整性、选择性和日志源对齐的语义判断上对保留的规则生成进行评分。Sigma-Hunter-Mistral总体得分为8.17,而最强通用基线为7.88,未微调的Mistral为4.61。有两个发现尤为突出:领域自适应使紧凑的7B模型能够在此结构化任务上与更大的通用模型竞争,而语法有效性是语义规则质量的弱代理,因为几个基线生成了格式良好的YAML但携带较弱的检测逻辑。自适应模型可在本地运行,这适合在分析师无法访问托管模型服务的隔离环境中进行检测工程。
英文摘要
Detection engineers must translate threat reports, forensic observations, and hunt hypotheses into precise, testable rules. General-purpose large language models (LLMs) can draft such rules, but often produce invalid YAML, incorrect log sources, unsupported fields, or overly broad detection logic. This paper presents \emph{Sigma-Hunter}, a domain-adapted LLM for analyst-assistive Sigma rule generation and threat hunting. We build an instruction-tuning dataset from 3,635 validated open-source Sigma rules, expanded into 7,663 question-answer and analyst-reasoning examples. Each source rule is assigned to a single train, validation, or test partition before this expansion, so no rule leaks across splits. We fine-tune a 7B Mistral model and a Phi-4 model with LoRA and score held-out rule generations on syntax, approximate field consistency, and a semantic judgment of detection logic, completeness, selectivity, and log-source alignment. Sigma-Hunter-Mistral scores 8.17 overall, against 7.88 for the strongest general-purpose baseline and 4.61 for untuned Mistral. Two findings stand out: domain adaptation enables a compact 7B model to perform competitively with larger general-purpose models on this structured task, and syntactic validity is a weak proxy for semantic rule quality, as several baselines emit well-formed YAML carrying weak detection logic. The adapted models run locally, which suits detection engineering in disconnected environments where analysts cannot reach hosted model services.