arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18496cs.CRcs.AI

MiST:面向网络安全的大语言模型中期训练

MiST: Mid-Training LLMs for Cybersecurity

Oded Ovadia, Elad Ben Zaken, Elad Guttman, Orly Moreno Kadosh

首次发表
浏览论文内容

中文总结 AI 辅助

MiST通过中期训练,利用专家审阅的种子语料生成合成数据,在网络安全基准上显著提升8B和32B模型性能,并为下游微调和强化学习提供更强初始化。

中文摘要 AI 辅助

网络安全结合了高风险分析与复杂的技术语言,使其成为大语言模型(LLM)一个具有影响力且充满挑战的领域。我们提出了MiST(中期训练的安全Transformer),这是一套8B和32B参数的模型,在公开的网络安全基准测试上取得了强劲性能。我们使用中期训练作为通用预训练与网络安全训练之间的一个中间适应阶段。我们不是对大量原始领域文本进行持续预训练,而是精心策划了一个紧凑的、经专家审阅的种子语料库,并将其转化为高质量的领域特定合成训练数据。最终的MiST检查点在平均网络安全准确率上,相对于对应的Qwen基线,在8B和32B模型上分别提升了+13.1和+8.6个绝对百分点,对应相对增益分别为+27.0%和+15.8%。消融实验进一步表明,这些网络安全性能的提升源于中期训练和监督微调阶段中合成数据生成流程的组合。此外,我们展示了MiST为下游任务特定的微调适应和强化学习提供了更强的初始化。

英文摘要

Cybersecurity combines high-stakes analysis with complex technical language, making it an impactful and challenging domain for LLMs. We present MiST (Mid-trained Security Transformer), a suite of 8B and 32B models that achieve strong performance on public cybersecurity benchmarks. We use mid-training as an intermediate adaptation stage between general pre-training and cybersecurity training. Rather than performing continual pre-training over large volumes of raw domain text, we curate a compact, expert-vetted seed corpus, and transform it into high-quality domain-specific synthetic training data. The final MiST checkpoints improve mean cybersecurity accuracy by +13.1 and +8.6 absolute percentage points over the corresponding Qwen baselines for 8B and 32B, respectively, corresponding to relative gains of +27.0% and +15.8%. Ablation results further show that these cybersecurity gains arise in the mid-training and supervised fine-tuning stages through a combination of the synthetic data generation flows. Furthermore, we show that MiST provides a stronger initialization for downstream task-specific fine-tuning adaptation and reinforcement learning.

发表机构

  • Dream(梦想)

机构由 AI 辅助整理,请以论文原文为准。

↑