AI 中文总结
CoSA是一种基于大语言模型的上下文感知漏洞严重程度评估方法,通过两阶段仓库剪枝策略与Transformer预测器,在6816个CVSS标注实例上较最优基线提升了14.4%准确率与15.3% Macro-F1
AI 中文摘要
准确的漏洞严重程度评估对于优先处理修复工作至关重要,但手动评估通用漏洞评分系统(CVSS)基础指标仍十分耗时。现有自动化方法往往无法捕捉评估多项CVSS基础指标所需的仓库级证据,而仓库感知评估颇具挑战性,因为相关证据分散在整个仓库中且存在大量噪声。为应对这些挑战,本文提出CoSA,一种基于上下文的漏洞严重程度评估方法,可从仓库构件中推断CVSS基础指标。CoSA构建代码属性图(CPG)并采用两阶段仓库剪枝策略:轻量级静态剪枝以保留结构邻近上下文,随后通过智能体式大语言模型(LLM)引导的剪枝步骤保留与CVSS相关的上下文并收集支撑证据。LLM随后将检索到的仓库上下文整合为紧凑的、按CVSS指标划分的文本摘要,输入轻量级Transformer预测器。本文还构建了一个更高质量的仓库级数据集,包含6816个标注有CVSS的实例,覆盖90种通用弱点枚举(CWE)类型。对真实漏洞的实验表明,CoSA的性能始终优于函数级基线和纯LLM基线,相比表现最佳的基线,它将预测准确率提升了14.4%,Macro-F1提升了15.3%,这表明面向指标的显式仓库上下文检索对于实用且可靠的自动化严重程度评估至关重要。
英文摘要
Accurate vulnerability severity assessment is essential for prioritizing remediation, yet manually assessing Common Vulnerability Scoring System (CVSS) base metrics remains labor-intensive. Existing automated approaches often fail to capture the repository-level evidence required for assessing many CVSS base metrics. Such repository-aware assessment is challenging because relevant evidence is scattered across the entire repository under heavy noise. To address these challenges, we present CoSA, a Context-aware vulnerability Severity Assessment approach that infers CVSS base metrics from repository artifacts. CoSA constructs a code property graph (CPG) and applies a two-stage repository-pruning strategy: lightweight static pruning to preserve structurally proximal context, followed by an agentic large language model (LLM)-guided pruning step to retain CVSS-relevant context while collecting supporting evidence. The LLM then consolidates the retrieved repository context into compact, CVSS metric-wise textual summaries, which are fed into a lightweight transformer predictor. We also construct a higher-quality repository-level dataset comprising 6,816 CVSS labeled instances spanning 90 Common Weakness Enumeration (CWE) types. Experiments on real-world vulnerabilities show that CoSA consistently outperforms function-level and pure-LLM baselines. It improves prediction accuracy by 14.4% and Macro-F1 by 15.3% over the best-performing baseline, suggesting that explicit, metric-oriented repository context retrieval is crucial for practical and reliable automated severity assessment.