arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ARQ:用于C/C++漏洞检测的智能CodeQL查询优化框架

ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection

Chunyi Wang, Yunfei Ke, Junfeng Yang, Yun-Yun Tsai, Penghui Li

arXiv 2608.20637首次发表:更新:

发表机构

Columbia University(哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ARQ智能框架,利用合成C/C++程序的执行证据和LLM优化CodeQL查询,无需标注数据等,优化后查询的真阳性最多提升119.8%,还修复了长期悬而未决的问题并发现新漏洞。

AI 中文摘要

静态分析工具已被广泛应用于C/C++程序的漏洞检测,基于查询的静态分析工具(如CodeQL)将易受攻击的代码模式编码为检测查询,并与源代码进行匹配。然而,现有的查询仍存在误报(FPs,将良性代码错误标记为易受攻击)和漏报(FNs,遗漏真实漏洞)问题。本文提出ARQ,一种智能框架,可利用从合成C/C++程序中获取的执行证据自动优化C/C++ CodeQL查询。核心思路是:当合成程序的执行结果与查询的判定不一致时,该程序会暴露查询的缺陷;若程序确实存在漏洞但查询未触发,则查询存在漏报缺陷;若程序安全但查询触发,则存在误报缺陷。ARQ随后运行基于大语言模型(LLM)的优化循环,以这些不一致结果为真实值修复查询。与以往的查询优化方法不同,ARQ无需标注数据集、提交历史或特定漏洞模板。本文通过使用三个商用大语言模型(GPT-5.4、Claude-Sonnet-4.6和Gemini-3.5-flash)优化12个官方CodeQL查询来验证ARQ的有效性,在Juliet v1.3和FormAI v2数据集上对比ARQ优化后的查询与原始CodeQL查询,结果显示ARQ优化后的查询检测到的真阳性数量显著增加,最高提升119.8%,且全程准确率至少为98.0%。ARQ成功修复了官方CodeQL查询仓库中三个未解决的GitHub问题,这些问题已悬而未决长达27个月,优化后的查询还在真实库libpng和zlib中发现了两个此前未被发现的漏洞。

英文摘要

Static analyzers have been widely adopted for vulnerability detection in C/C++ programs. Query-based static analyzers (e.g., CodeQL) encode vulnerable code patterns in detection queries and match them against source code. However, existing queries still suffer from false positives (FPs, incorrectly flagging benign code as vulnerable) and false negatives (FNs, missing real vulnerabilities). We present ARQ, an agentic framework that automatically refines C/C++ CodeQL queries using execution-grounded evidence from synthesized C/C++ programs. Our key insight is that a synthesized program exposes a query's weakness whenever its execution disagrees with the query's verdict. If the program is genuinely vulnerable but the query stays silent, the query has an FN weakness; if the program is safe but the query fires anyway, it has an FP weakness. ARQ then runs an LLM-based refinement loop that repairs the query using these disagreements as ground truth. Unlike previous query refining methods, ARQ requires no labeled datasets, no commit history, and no vulnerability-specific templates. We demonstrate the effectiveness of ARQ by refining 12 official CodeQL queries using three commercial LLMs (GPT-5.4, Claude-Sonnet-4.6, and Gemini-3.5-flash). We compare both ARQ-refined and original CodeQL queries on the Juliet v1.3 and FormAI v2 datasets and show that ARQ-refined queries detect substantially more true positives, by up to 119.8\%, with a Precision of at least 98.0\% throughout. ARQ successfully fixed three unresolved GitHub issues raised in the official CodeQL query repository that had remained open for as long as \textit{27 months}. The refined queries also exposed two previously undiscovered bugs in the real-world libraries libpng and zlib.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑