arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Java漏洞的CodeQL误报与查询优化的实证分析

An Empirical Analysis of CodeQL False Positives and Query Refinements for Java Vulnerabilities

Amirali Sajadi, Saikat Dutta, Preetha Chatterjee

arXiv 2609.04535首次发表:更新:

发表机构

College of Computing & Informatics Drexel University; Department of Computer Science Cornell University(德雷塞尔大学计算与信息学院; 康奈尔大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究分析Java漏洞的CodeQL误报模式,实现查询优化可减少大量误报,还验证了智能编码工具能将优化模式适配新项目,支持以优化为导向的SAST工作流以降低排查工作量。

AI 中文摘要

静态应用安全测试(SAST)工具可帮助开发者在部署前发现漏洞,但误报会产生大量分类排查工作。本研究探讨Java安全分析中的CodeQL误报是否会形成可通过优化分析来减少的、可重复且可解释的模式。我们针对来自110个项目的167个CVE实例运行CodeQL的Java安全查询套件,重点关注误报率最高的10个查询。我们手动审查了500个抽样的误报路径与位置,并构建了源代码层面的分类体系,该分类包含五类:路径约束或净化措施遗漏(占36.6%)、良性执行上下文(占29.4%)、信任边界建模缺失(占27.6%)、并发建模不精确(占5%)以及 sink 建模不精确(占1.4%)。基于这些发现,我们实现了CodeQL查询优化,可在查询层面检测并过滤重复出现的误报模式。这些优化移除了81.8%的被审查误报;在全部选定查询数据集上,它们移除了15.8%的报告路径与位置,同时保留了8个真实漏洞中的7个。这表明许多误报可在分析阶段减少,尽管固定的优化常依赖于特定项目的上下文。为解决这种泛化差距,我们评估智能编码工具是否可将优化模式适配到新项目。以我们的模式为模板,两款工具分别在56%和62%的任务上取得成功,且查询编译通过率超过90%;若无此指导,两款工具的成功率仅为28%,编译率降至30%-36%。这些结果支持以优化为导向的SAST工作流,即通过在CodeQL查询中对重复误报建模并自动适配不同项目上下文,减少重复分类排查工作。

英文摘要

Static application security testing (SAST) tools help developers find vulnerabilities before deployment, but false positives create substantial triage effort. We study whether CodeQL false positives in Java security analysis form recurring, explainable patterns that can be reduced by refining the analysis. We run CodeQL's Java security query suite on 167 CVE instances from 110 projects, focusing on the ten queries with the highest false positive rates. We manually review 500 sampled false positive paths and locations and construct a source-level taxonomy. The five categories are Missed Path Constraint or Sanitization (36.6%), Benign Execution Context (29.4%), Missing Trust Boundary Modeling (27.6%), Imprecise Concurrency Modeling (5%), and Imprecise Sink Modeling (1.4%). Guided by these findings, we implement CodeQL refinements that detect and filter recurring false positive patterns at the query level. The refinements remove 81.8% of reviewed false positives. Across the full selected-query dataset, they remove 15.8% of reported paths and locations while retaining 7 of 8 true positives. This shows that many false positives can be reduced in the analysis, although fixed refinements often depend on project-specific context. To address this generalization gap, we evaluate whether agentic coding tools can adapt refinement patterns to new projects. Given our patterns as templates, the two tools succeed on 56% and 62% of tasks, with query compile-pass rates above 90%. Without this guidance, both succeed on only 28%, while compile rates fall to 30-36%. These results support a refinement-oriented SAST workflow in which recurring false positives are modeled in CodeQL queries and automatically adapted to different project contexts, reducing repeated triage.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑