利用上下文学习预测漏洞严重性:一项工业案例研究
On Predicting Vulnerability Severity Using In-Context Learning: An Industrial Case Study
浏览论文内容
中文总结 AI 辅助
本研究通过工业案例,用可本地部署的开源LLM的上下文学习从C/C++代码片段预测CVSS v3.1评分,发现CodeLlama2-7B在轻量输出约束提示下可接近云服务性能,且Big-Vul可作为工业数据代理。
中文摘要 AI 辅助
现代软件系统需要更早、更具可扩展性的漏洞严重性评估,以减少高影响安全漏洞的暴露。安全分析师通常会分配CVSS评分,但这种手动分类无法跟上公开漏洞的增长,且往往依赖云LLM服务,引发保密性问题。本文开展工业案例研究,利用可本地部署的开源LLM的上下文学习,直接从易受攻击的C/C++代码片段预测CVSS v3.1评分。我们将专有数据与Big-Vul数据集对比,显示CVSS分布足够一致,证明在构建基于提示的测试平台时,Big-Vul可作为工业数据的代理。随后我们调整上下文配置和模型参数,使用均方误差(MSE)和可行性指标评估CodeLlama2-7B、CodeLlama2-13B、Mistral-7B、gpt-oss和GPT4o-mini。结果表明,中等规模的开源代码模型,尤其是CodeLlama2-7B,在轻量、输出约束提示的引导下,可接近云服务在CVSS回归任务上的最佳性能,为工业场景的严重性分类提供了实用、隐私保护的基础组件。
英文摘要
Modern software systems require earlier and more scalable vulnerability severity assessment to reduce exposure to high-impact security flaws. Security analysts typically assign CVSS scores, but this manual triage does not scale with the growth of disclosed vulnerabilities and often depends on cloud LLM services that raise confidentiality concerns. This paper presents an industrial case study on predicting CVSS v3.1 scores directly from vulnerable C/C++ snippets using in-context learning with locally deployable, open-source LLMs. We compare proprietary data with the Big-Vul dataset, showing sufficiently aligned CVSS distributions to justify Big-Vul as a proxy for industrial data when constructing prompt-based testbeds. We then vary in-context configurations and model parameters, evaluating CodeLlama2-7B, CodeLlama2-13B, Mistral-7B, gpt-oss, and GPT4o-mini using mean squared error (MSE) and feasibility metrics. Our results show that medium-sized open-source code models, particularly CodeLlama2-7B, can approximate the best cloud performance for CVSS regression when guided by lightweight, output-constraining prompts, offering a practical, privacy-preserving building block for severity triage in industrial settings.
发表机构
- William & Mary(威廉玛丽学院)
- Microsoft(微软公司)
- Cisco Systems(思科系统公司)
机构由 AI 辅助整理,请以论文原文为准。