使用具有错误减少技术的大语言模型来判定静态分析警报
Using LLMs to Adjudicate Static-Analysis Alerts with Error Reduction Techniques
AI总结:
研究用大语言模型判定静态分析警报,采用一致性检查和LLM推理评估两种方法,在三个测试套件上评估多个LLMs,中级推理LLMs经错误缓解后召回率和特异性高,还探讨了相关问题及合成程序作为证据的结果。
AI中文摘要:
静态分析广泛用于在部署前查找源代码中的安全弱点,但它产生的警报远超分析师可审查的数量。我们研究大语言模型(LLMs)判定静态分析警报(分类为真实漏洞或误报)的能力。使用两种错误缓解方法:(1)一致性检查(CC),多次运行LLM并检查判定是否一致;(2)LLM推理评估(LRE)步骤,多次运行LLM后让其根据每次运行提供的推理选择判定。在三个测试套件Juliet、FormAI和SV-COMP上评估了多个LLMs。测试的中级推理LLMs在所有三个套件中都达到了高召回率和特异性。通过错误缓解,它们在每个套件上至少达到98%的召回率和至少94.8%的特异性。还探讨了Juliet记忆问题,基于FormAI得出泛化结论。此外报告了使用LLM合成动态触发漏洞程序作为独立证据的结果,有效触发被证明是真实漏洞的有力证据。
英文摘要:
Static analysis is widely used for finding security weaknesses in source code before deployment, but it often produces far more alerts than analysts can review. We study how well large language models (LLMs) can adjudicate (classify as a real bug or a false alarm) static-analysis alerts. We use two mistake-mitigation methods: (1) a consistency check (CC) that runs the LLM multiple times and checks that the verdicts are consistent with each other, and (2) an LLM reasoning evaluation (LRE) step that runs the LLM multiple times and then asks the LLM to choose a verdict after evaluating the reasoning provided by each run. We evaluated several LLMs on three test suites: Juliet, FormAI, and SV-COMP. Across all three suites, the mid-tier reasoning LLMs that we tested (o4-mini, gpt-oss-120b, gpt-oss-20b) reach high recall (percent of real bugs that the tool correctly flags as needing repair / manual attention) and specificity (percent of actually false alerts that the tool correctly dismisses as false alarms). With mistake mitigation, they reach at least 98% recall and at least 94.8% specificity on every suite (with CC alone on Juliet and SV-COMP, and with LRE+CC on FormAI). We probe Juliet memorization and show that o4-mini can often reconstruct sanitized test cases' original identities, so we base our generalization claims primarily on FormAI, scored against our own unpublished manual adjudications. We also report results of using the LLM to synthesize a program that dynamically triggers the flaw as independent evidence; a validity check rejected every trigger driver aimed at a false alarm, so a valid trigger proved to be strong evidence of a real flaw.