arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28460cs.LGcs.CR

基于推理增强语言模型的网络安全检测分类

Cybersecurity Detection Classification with Reasoning-enabled Language Models

Amol Khanna, Manu Nandan, Cristian Viorel Popa, Joan Pujol-Roig, Diana Bolocan, Laura Vasilie, Alexandru Apostu, Chase Helwig, Mihaela Gaman, Michael Brautbar, … 展开作者

Amol Khanna, Manu Nandan, Cristian Viorel Popa, Joan Pujol-Roig, Diana Bolocan, Laura Vasilie, Alexandru Apostu, Chase Helwig, Mihaela Gaman, Michael Brautbar, Edward Raff, Chase Midler, Sven Krasser

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对SOC警报疲劳问题,训练了思维链推理增强的分类器及校准器,在Windows终端检测任务上提升了良性与恶意召回率,且30B微调模型优于通用大模型。

中文摘要 AI 辅助

安全运营中心(SOC)的一个主要问题是警报疲劳,因为每天报告的检测数量超过了工作人员能够分类的数量。现有工作通过提示或微调大型语言模型(LLM)直接输出分类标签,但未训练它们推理检测是否为真正威胁。我们结合自动提示优化、自训练和带可验证奖励的强化学习,在真实人工标注的Windows终端检测数据上训练了一个思维链(CoT)推理增强的分类器。我们发现CoT推理会降低自动分类所依赖的标签token概率,因此我们单独训练了一个校准器,该校准器读取完整推理轨迹并估计判定正确的概率。我们的系统达到82.6%的测试准确率,在控制自动分类的高置信度操作点,与直接标签LLM分类器相比,良性召回率提高43.0%,恶意召回率提高18.3%。我们进一步表明,训练后的校准器是必要的——未训练的置信度判断会使高置信度召回率降至零,且微调后的30B模型明显优于前沿通用模型,这推动了针对规模的针对性训练。

英文摘要

A major issue in Security Operations Centers (SOCs) is alert fatigue, as the number of detections reported is more than staff can triage in a given day. Prior work prompts or fine-tunes large language models (LLMs) to emit a triage label directly, but does not train them to reason about whether a detection is a genuine threat. We train a chain-of-thought (CoT) reasoning-enabled triage classifier on real, human-labeled Windows endpoint detections by combining automated prompt optimization, self-training, and reinforcement learning with verifiable rewards. We find that CoT reasoning also degrades the label-token probabilities that automated triage relies on, so we separately train a calibrator that reads the full reasoning trace and estimates the probability that the verdict is correct. Our system reaches 82.6% test accuracy and, at the high-confidence operating point that governs automated triage, improves benign recall by 43.0% and malicious recall by 18.3% over a direct-label LLM classifier. We further show that the trained calibrator is necessary - an untrained confidence judge collapses high-confidence recall to zero - and that a finetuned 30B model significantly outperforms frontier general-purpose models, motivating targeted training over scale.

↑