面向可扩展且成本高效的漏洞检测:自动查询生成研究
Towards Scalable and Cost-Efficient Vulnerability Detection: A Study on Automatic Query Generation
浏览论文内容
中文总结 AI 辅助
本研究利用国家漏洞数据库数据,实证评估大型语言模型自动生成CodeQL查询的能力,结果显示其显著提升基线查询,平均F1分数提高82%,并证明该方法是大规模漏洞检测中可扩展且成本高效的替代方案。
中文摘要 AI 辅助
静态分析仍然是软件安全的基石,然而诸如 CodeQL 等工具的有效性往往受限于开发高覆盖率查询套件所需的大量人工努力。尽管大型语言模型(LLM)已作为自动化代码推理的潜在解决方案出现,但其在生成结构化、可执行安全查询方面的实际效用仍未得到充分探索。在本文中,我们开展了一项实证研究,利用国家漏洞数据库中的漏洞数据来评估 LLM 综合生成 CodeQL 查询的能力。通过这一调查,我们探索了将 LLM 用作自动 CodeQL 查询生成器的潜力。随后,我们系统地评估了各种 LLM 架构在一组多样化真实世界漏洞上的性能,衡量它们提高检测覆盖率和精度的能力。我们的研究结果表明,LLM 生成的查询显著增强了基线 CodeQL 查询,平均 F1 分数提升了 82%。此外,我们提供了详细的成本效益分析,显示虽然直接基于 LLM 对整个代码库进行扫描通常在计算和财务上代价高昂,但利用 LLM 综合生成 CodeQL 查询为大规模漏洞检测提供了一种可扩展且成本高效的替代方案。我们的结果建议,LLM 能有效弥合非结构化漏洞报告与正式静态分析规范之间的差距,为全面自动化漏洞检测提供一条可扩展的路径。
英文摘要
Static analysis remains a cornerstone of software security, yet the effectiveness of tools such as CodeQL is often limited by the substantial manual effort required to develop high-coverage query suites. While large language models (LLMs) have emerged as a potential solution for automated code reasoning, their practical utility in generating structured, executable security queries remains underexplored. In this paper, we conduct an empirical study to evaluate the ability of LLMs to synthesize CodeQL queries using vulnerability data from the National Vulnerability Database. Through this investigation, we explore the potential of using LLMs as an automatic CodeQL query generator. Subsequently, we systematically evaluate the performance of various LLM architectures across a diverse set of real-world vulnerabilities, measuring their ability to improve detection coverage and precision. Our findings reveal that LLM-generated queries significantly enhance the baseline CodeQL queries, yielding 82% improvement in average F1-score. Furthermore, we provide a detailed cost- benefit analysis showing that while direct LLM-based scanning of entire repositories is often computationally and financially prohibitive, leveraging LLMs to synthesize CodeQL queries offers a scalable and cost-effective alternative for large-scale vulnerability detection. Our results suggest that LLMs can effectively bridge the gap between unstructured vulnerability reports and formal static analysis specifications, offering a scalable path toward comprehensive automated vulnerability detection.
发表机构
- Singapore Management University(新加坡管理大学)
- Monash University(蒙纳士大学)
- University of Sydney(悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。