AI 中文总结
本文对比分析了逻辑回归、SetFit等多种安全漏洞报告自动识别技术,发现SetFit整体性能最优,GPT-5.2表现较差,迁移学习对不同数据量项目的性能影响存在差异。
AI 中文摘要
及时识别与安全相关的漏洞报告,对于缩小软件系统的漏洞窗口至关重要。手动筛查收到的漏洞报告以识别安全相关问题,对于大型软件系统而言耗时、易出错且无法扩展。因此,已有多种自动技术被提出以推动该任务,包括传统机器学习(ML)技术和大语言模型。但现有文献仍较为零散:多数研究仅引入或优化某一特定技术,并在有限的基准集及不同实验设置下对其进行评估,导致难以对比各研究结果、得出关于现有方法有效性的可靠结论,也无法为研究人员和从业者明确哪些技术最适合该任务。为填补这一空白,本文使用基准数据集对多种有前景的安全漏洞报告自动识别技术开展对比分析,评估的方法包括逻辑回归、支持向量机、随机森林、OpenAI的GPT-5.2、BERT-base、RoBERTa,以及当前最优的少样本学习框架SetFit。结果显示,SetFit整体性能最优,F1值达0.80,在四个数据集中的三个上表现优于其他技术;RoBERTa表现具竞争力,在部分项目中接近SetFit;传统ML技术(尤其是逻辑回归)在特定场景下仍是强基准。相反,GPT-5.2在零样本和少样本设置下表现均较差。此外,跨项目实验表明,迁移学习可提升数据有限项目的性能,但可能降低具有强项目特定特征项目的结果。
英文摘要
Timely identification of security-related bug reports is essential to minimize the window of vulnerabilities in software systems. Manually screening incoming bug reports to identify security-related issues is time-consuming, error-prone, and non-scalable for large-scale software systems. Thus, a variety of automatic techniques, including traditional machine learning (ML) techniques and large language models, have been proposed to facilitate this task. However, the literature remains fragmented. Most studies introduce or optimize a particular technique and evaluate it against a limited set of baselines, often under different experimental setups. As a result, it is difficult to compare their results and draw reliable conclusions about the effectiveness of existing approaches, leaving researchers and practitioners without clear guidance on which techniques are most suitable for the task. To address this gap, we conducted a comparative analysis of several promising automated techniques to identify security-related bug reports using benchmark datasets. We evaluated Logistic Regression, Support Vector Machines, Random Forest, OpenAI's GPT-5.2, BERT-base, RoBERTa, and SetFit (a state-of-the-art few-shot learning framework). Our results indicate that SetFit achieves the best overall performance, achieving an F1-score of 0.80 and outperforming other techniques on three of the four datasets. RoBERTa performs competitively and approaches SetFit in some projects, while traditional ML techniques, particularly Logistic Regression, remain a strong baseline in certain contexts. In contrast, GPT-5.2 performs poorly in both zero-shot and few-shot settings. In addition, cross-project experiments demonstrate that transfer learning can improve performance for projects with limited data, but may degrade results for projects with strong project-specific characteristics.